Intelligent shopping decision-making method and device based on voice recognition, equipment and medium

Through voice enhancement, voiceprint feature extraction and phoneme decoding, combined with the consumption intention matrix, the problems of low efficiency and insufficient accuracy of speech recognition in the prior art are solved, and more efficient and accurate consumption decisions and recommendations are achieved.

CN120279891APending Publication Date: 2025-07-08PING AN HEALTH INSURANCE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510548249.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing speech recognition technology has insufficient semantic understanding and full-process coverage capabilities in the fields of medical health and financial technology, resulting in low efficiency in drug purchases, poor accuracy in health insurance recommendations, and prone to misjudgment of intentions and process interruptions under high concurrent pressure.

Method used

Through voice enhancement, voiceprint feature extraction, phoneme decoding and feature fusion, a consumption intention matrix is built to achieve the consumption decisions of target users.

Benefits of technology

It improves the accuracy and robustness of speech recognition, enhances the accuracy and efficiency of consumer decisions, can more accurately capture user needs, and recommends products or services that meet personalized conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279891A_ABST
    Figure CN120279891A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice semantics, can be applied to business system platforms of financial science and technology, medical health and the like, and discloses an intelligent shopping decision-making method, device, equipment and medium based on voice recognition. Obtaining target voice data; performing voiceprint feature extraction on the target voice data to obtain a target voiceprint feature vector; performing phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector; performing feature fusion according to the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector; and constructing a consumption intention matrix of the target user according to the comprehensive semantic vector, and performing consumption decision according to the consumption intention matrix to obtain an optimal consumption scheme. According to the invention, the language recognition efficiency can be improved and full-process voice driving can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech semantics, and particularly to an intelligent shopping decision-making method, device, equipment and medium based on speech recognition. Background Art

[0002] In the fields of medical health and fintech, application scenarios of online shopping experiences based on speech recognition technology are gradually increasing, but there are still many deficiencies in the existing technology in terms of the depth of semantic understanding and the ability to cover the entire process.

[0003] For example, in the intelligent selection of drugs in the field of medical health, most of the existing technology users can only perform voice queries and purchases of drugs by drug names or specific keywords. In such a large-scale drug library, it is very difficult to remember the drug names, drug effects, dosages and other drug information of each drug. Using this method of querying and purchasing based on drug names or keywords, the efficiency of prescribing drugs is low, and the pertinence and accuracy of prescribing drugs are also difficult to guarantee.

[0004] For example, in the field of fintech business, the existing speech recognition technology has insufficient semantic understanding accuracy in the recommendation and purchase of health insurance products. For example, health insurance involves professional terms such as "pre-existing condition exemption" and "waiting period". Traditional speech recognition models lack the support of a vertical domain knowledge graph and are prone to misjudgment of intentions for colloquial expressions such as "I have three highs and want to buy hospitalization insurance". At the same time, insurance decisions require continuous questioning of in-depth information such as family structure and medical treatment habits. The existing real-time dialogue state tracking technology is not mature, and the accuracy of context association decreases after more than multiple rounds of interaction, forcing users to repeat basic information, resulting in limitations in personalized recommendations.

[0005] The existing speech recognition technology is mainly based on an instruction-based interaction architecture. For example, basic functions are realized through a limited instruction set such as presetting "search for mobile phone" and "slide the commodity". Although this method can complete simple operations, the recognition accuracy of non-standard expressions such as elliptical sentences and inverted sentences in the natural language processing layer is insufficient; at the personalized adaptation level, it is difficult to dynamically combine the user's historical behavior to analyze fuzzy requests. Especially during the peak promotion period, the concurrent pressure and semantic complexity of voice requests increase exponentially, and the existing technology is prone to instruction misjudgment or process interruption. Therefore, how to improve the language recognition efficiency and achieve full-process voice drive has become an urgent problem to be solved. Summary of the Invention

[0006] The present invention provides an intelligent shopping decision-making method, device, equipment and medium based on speech recognition, and its main purpose is to solve the problem of low efficiency of voice interaction.

[0007] In a first aspect, to achieve the above object, an intelligent shopping decision-making method based on speech recognition provided by the present invention includes:

[0008] Obtain the voice shopping data of the target user, perform voice enhancement on the voice shopping data to obtain target voice data;

[0009] Extract the voiceprint feature from the target voice data to obtain a target voiceprint feature vector;

[0010] Perform phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector;

[0011] Perform feature fusion based on the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector;

[0012] Construct a consumption intention matrix of the target user according to the comprehensive semantic vector, and perform consumption decision-making according to the consumption intention matrix to obtain an optimal consumption plan.

[0013] In a second aspect, the present invention also provides an intelligent shopping decision-making device for voice recognition, including:

[0014] A voice enhancement module, configured to obtain the voice shopping data of the target user, perform voice enhancement on the voice shopping data to obtain target voice data;

[0015] A feature extraction module, configured to extract the voiceprint feature from the target voice data to obtain a target voiceprint feature vector;

[0016] A phoneme decoding module, configured to perform phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector;

[0017] A feature fusion module, configured to perform feature fusion based on the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector;

[0018] A consumption decision-making module, configured to construct a consumption intention matrix of the target user according to the comprehensive semantic vector, and perform consumption decision-making according to the consumption intention matrix to obtain an optimal consumption plan.

[0019] In a third aspect, the present invention also provides an electronic device, the electronic device includes:

[0020] At least one processor; and,

[0021] A memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the above-mentioned intelligent shopping decision-making method based on voice recognition.

[0023] Fourthly, the present invention also provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned intelligent shopping decision-making method based on speech recognition.

[0024] The present invention significantly improves the quality of speech signals through speech enhancement, can effectively reduce the influence of noise, makes the target speech clearer and more distinguishable, thereby improving the accuracy of speech recognition; gradually refines the voiceprint features through multi-level processing, from the original signal to audio features, then to the preliminary voiceprint representation, and finally obtains a highly discriminative voiceprint feature vector, improving the accuracy and robustness of voiceprint recognition; reduces the phoneme recognition error rate through four-level progressive processing, greatly improves the decoding accuracy, and realizes the lossless fusion of voiceprint and phoneme features through the dynamic dimension alignment algorithm, improving the cross-feature information fusion ability, and also improving the accuracy of speech data recognition; by constructing a consumption intention matrix, it can more accurately capture and analyze the consumption preferences and needs of target users, thereby recommending products or services that better meet their personalized needs for users, improving the consumption decision-making efficiency, and enhancing the accuracy of consumption decisions. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments of the present invention. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0026] Figure 1 It is a schematic diagram of an application environment of an intelligent shopping decision-making method based on speech recognition in an embodiment of the present invention;

[0027] Figure 2 It is a schematic flowchart of an intelligent shopping decision-making method based on speech recognition provided by an embodiment of the present invention;

[0028] Figure 3 It is a schematic flowchart of speech enhancement for the voice shopping data provided by an embodiment of the present invention;

[0029] Figure 4 It is a schematic flowchart of constructing the consumption intention matrix of the target user according to the comprehensive semantic vector provided by an embodiment of the present invention;

[0030] Figure 5 It is a schematic diagram of the modules of an intelligent shopping decision-making device based on speech recognition provided by an embodiment of the present invention;

[0031] Figure 6Schematic structural diagram of an electronic device for implementing an intelligent shopping decision-making method based on speech recognition according to an embodiment of the present invention;

[0032] Figure 7 Another schematic structural diagram of an electronic device for implementing an intelligent shopping decision-making method based on speech recognition according to an embodiment of the present invention.

[0033] The realization, functional features and advantages of the present invention will be further described with reference to the embodiments and the accompanying drawings. Specific embodiments

[0034] In order to enable those skilled in the art of the present technology to better understand the technical solutions of the present disclosure, and to fully understand how the present disclosure uses technical means to solve technical problems and the implementation process of achieving corresponding technical effects and implement accordingly, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all of the embodiments. The embodiments of the present disclosure and each feature in the embodiments can be combined with each other without conflict, and the formed technical solutions are all within the protection scope of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.

[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, device, product or equipment including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or equipment.

[0036] An embodiment of the present application provides an intelligent shopping decision-making method based on speech recognition. The execution subject of the intelligent shopping decision-making method based on speech recognition includes, but is not limited to, at least one of electronic devices such as a server, a terminal, etc. that can be configured to execute the device provided in the embodiment of the present application. In other words, the intelligent shopping decision-making method based on speech recognition can be executed by software or hardware installed on a terminal device or a server device. The server includes, but is not limited to: a single server, a server cluster, a cloud server, or a cloud server cluster, etc. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0037] An intelligent shopping decision-making method based on speech recognition according to the present invention can be applied in an application environment such as Figure 1 . Among them, the client communicates with the server through the network. The server can obtain target voice shopping data through the client. By performing speech enhancement on the voice shopping data, the quality of the voice signal is significantly improved, the influence of noise can be effectively reduced, making the target voice clearer and more distinguishable, thereby improving the accuracy of speech recognition; by performing multi-level processing to gradually refine the voiceprint features, from the original signal to audio features, then to the initial voiceprint representation, and finally obtaining a highly discriminative voiceprint feature vector, improving the accuracy and robustness of voiceprint recognition; through four-level progressive processing, the phoneme recognition error rate is reduced and the decoding accuracy is greatly improved. Through the dynamic dimension alignment algorithm, the lossless fusion of voiceprint and phoneme features is achieved, improving the cross-feature information fusion ability and also improving the accuracy of voice data recognition; by constructing a consumption intention matrix, the consumption preferences and needs of the target user can be captured and analyzed more accurately, so as to recommend products or services that better meet the personalized needs of the user, improving the consumption decision-making efficiency and enhancing the accuracy of consumption decisions. Finally, the optimal consumption plan is output and fed back to the client. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.

[0038] Referring to Figure 2 shown, it is a schematic flowchart of an intelligent shopping decision-making method based on speech recognition provided by an embodiment of the present invention. In this embodiment, the intelligent shopping decision-making method based on speech recognition includes:

[0039] S1. Obtain the voice shopping data of the target user, perform voice enhancement on the voice shopping data, and obtain target voice data.

[0040] In the embodiments of the present invention, the voice shopping data can record voice shopping data and interaction context through terminal devices such as smart speakers, in-vehicle voice systems, and mobile voice assistants. For example, when the user says "Buy a box of amoxicillin", the terminal device will store the audio file, timestamp, device ID, and the converted text instruction, and associate the shopping cart operation record of the target user account (such as whether to modify the quantity or cancel the order). The voice activity detection (VAD) technology is used to filter out invalid segments and retain the valid shopping-related voice.

[0041] Exemplarily, in the fintech business scenario, in order to better meet the needs of users, financial institutions such as banks and insurance companies use smart voice assistants to handle various financial services such as financial product purchases and insurance claims. When users need to handle financial services provided by financial institutions such as purchasing financial products and insurance claims, they can send voice instructions or voice questions to the voice assistant through terminal devices such as smartphones, smart speakers, and computers. After the voice assistant performs business handling actions such as purchase actions and claim actions according to the voice instructions or voice questions, it feeds back the results of business handling to the user in a voice manner.

[0042] As Figure 3 shown, in the embodiments of the present invention, the performing voice enhancement on the voice shopping data to obtain target voice data includes:

[0043] Decompose the voice shopping data to obtain noise spectrum features and user voice spectrum features;

[0044] Construct an adaptive filter according to the noise spectrum features and the parameters of the adaptively acquired adaptive filter;

[0045] Suppress noise from the user voice spectrum features according to the adaptive filter to obtain target voice data.

[0046] Specifically, the noise spectrum features mainly describe the characteristics of the noise signal in the frequency domain, including the frequency distribution and intensity of the noise, and the user voice spectrum features mainly describe the characteristics of the target user voice signal in the frequency domain, including the fundamental frequency and formants of the voice.

[0047] Specifically, the voice shopping data has non-stationary characteristics in the time domain (such as the start and end of the voice and pitch changes), and it is difficult to distinguish noise from valid voice components through direct analysis. The continuous voice data can be framed into segments with a length of 20 - 40 ms, and each framed voice data segment is subjected to a short-time Fourier transform to convert the time-domain waveform into a frequency-domain energy distribution, obtaining a time-frequency matrix.

[0048] Among them, the time-frequency matrix is decomposed to obtain a noise spectrum feature and a user speech spectrum feature. The adaptive filter can be a Wiener filter. Frequency-domain filtering is performed on each frequency point of the user speech spectrum according to the Wiener filter to remove or weaken the noise components therein, thereby obtaining target speech data.

[0049] Optionally, before enhancing the voice of the voice shopping data to obtain target voice data, the method further includes:

[0050] Using a deep learning algorithm to classify the noise spectrum features and separating the classification results according to different types of noise;

[0051] Dynamically adjusting the parameters of the adaptive filter according to the classification results to optimize the voice enhancement effect;

[0052] Real-time evaluating the voice enhancement effect and further adjusting the parameters of the adaptive filter according to the evaluation results.

[0053] Specifically, a convolutional neural network (CNN) can be used to extract and classify the features of the noise spectrum collected in real time. The short-time Fourier transform is used to convert the time-domain signal into a low-dimensional Mel spectrogram as the input. The time-frequency domain local features are extracted through three convolutional layers, and finally the probability distributions of six types of noise such as Gaussian noise, wind noise, and mechanical shock are output through the fully connected layer.

[0054] Among them, the classification can adopt a transfer learning strategy. After pre-training on the ESC-50 environmental sound data set, the algorithm weights are continuously updated through an online learning mechanism to improve the recognition ability of unknown noise.

[0055] Specifically, the normalized least mean square (NLMS) algorithm is adopted for steady-state noise and a small step size (μ = 0.01) is set to ensure stable convergence. For transient impact noise, the affine projection algorithm is switched and the step size is increased to 0.1 to improve the tracking speed; in wideband interference scenarios such as wind noise, the sub-band filtering structure is automatically enabled and the convergence factor is inversely correlated with the band energy. The parameter mapping library presets 32 typical noise-parameter combinations, and the fuzzy logic decision module is used to process the mixed noise situation in the classification probability distribution, and the weighted average method is used to generate the transition parameter configuration.

[0056] Meanwhile, a real-time two-dimensional evaluation mechanism can be innovatively introduced. In the objective dimension, the improved Perceptual Evaluation of Speech Quality (PESQ) and Short-Time Objective Intelligibility (STOI) metrics of the processed signal are calculated, and the evaluation results are updated every 200 ms. In the subjective dimension, a lightweight generative adversarial network is used to simulate the human ear's auditory perception, and a speech naturalness score ranging from 0 to 1 is output through the discriminator network. When the comprehensive score is lower than the threshold of 0.85, a parameter optimizer is triggered to perform a directed search within the preset parameter space based on the Bayesian optimization algorithm, and the configuration scheme with performance improvement is retained in the knowledge base after each adjustment.

[0057] In the embodiment of the present invention, the decomposing the voice shopping data into a noise spectral feature and a user voice spectral feature includes:

[0058] Performing signal conversion on the voice shopping data according to the short-time Fourier transform to obtain a time-frequency domain signal;

[0059] Extracting a noise spectral basis matrix and a voice spectral basis matrix from the time-frequency domain signal;

[0060] Performing feature extraction on the noise spectral basis matrix and the voice spectral basis matrix respectively to obtain a noise spectral feature and a user voice spectral feature.

[0061] Specifically, the noise spectral basis matrix refers to extracting the spectral basis matrix of the noise by analyzing the background noise in the time-frequency domain signal using matrix decomposition techniques (such as non-negative matrix factorization (NMF), singular value decomposition (SVD), etc.), and the noise spectral basis matrix describes the distribution characteristics of the noise in the frequency domain.

[0062] Among them, a matrix decomposition technique can be used to extract the spectral basis matrix of the voice from the time-frequency domain signal, and the voice spectral basis matrix describes the main characteristics of the voice shopping data in the frequency domain, such as the fundamental frequency, formants, etc.

[0063] In the embodiment of the present invention, voice enhancement significantly improves the quality of the voice signal. In a shopping scenario, the target user may be in a noisy environment, such as a public place, a vehicle, etc., and the background noise will seriously interfere with the accuracy of speech recognition. Through voice enhancement, the influence of noise can be effectively reduced, making the target voice clearer and more distinguishable, thereby improving the accuracy of speech recognition.

[0064] In the embodiment of the present invention, the decomposing the voice shopping data into a noise spectral feature and a user voice spectral feature, the method further includes:

[0065] Using a deep learning model (such as a convolutional neural network) to automatically learn spectral features to improve the accuracy and robustness of feature extraction;

[0066] Introduce an attention mechanism to enable the deep learning model to pay more attention to the key features in the speech spectrum, further improving the speech enhancement effect.

[0067] Specifically, use a deep learning model (such as a convolutional neural network) to automatically learn spectral features. Decompose the original time-domain signal into intrinsic mode functions through improved variational mode decomposition, and combine adaptive wavelet packet transform to perform secondary separation on the high-frequency noise components to construct a time-frequency matrix. Process the noise and speech spectral features respectively through a two-channel deep convolutional network to enhance the ability to analyze speech formants.

[0068] S2. Extract the voiceprint features from the target speech data to obtain a target voiceprint feature vector.

[0069] In the embodiments of the present invention, the target voiceprint feature vector has relative stability and can be used as the representation and identification of the target speaker, which is a voice feature that distinguishes the current target speaker from other speakers. In order to distinguish the current speaker from the simulated speaker and ensure the authenticity of the identity of the current speaker, voiceprint features are extracted from the target speech data to extract effective, stable, and reliable features that uniquely represent the identity of the target speaker.

[0070] In the embodiments of the present invention, the extracting the voiceprint features from the target speech data to obtain a target voiceprint feature vector includes:

[0071] Extract audio signal features from the target speech data to obtain audio signal features;

[0072] Perform first voiceprint feature extraction on the audio signal features to obtain preliminary voiceprint features;

[0073] Perform second voiceprint feature extraction on the preliminary voiceprint features to obtain a target voiceprint feature vector.

[0074] In the embodiments of the present invention, the extracting the voiceprint features from the target speech data to obtain a target voiceprint feature vector further includes: using multi-modal data (such as facial expressions or gestures) to assist in the extraction of voiceprint features to enhance the uniqueness and stability of the features;

[0075] Further, after performing second voiceprint feature extraction on the preliminary voiceprint features to obtain a target voiceprint feature vector, it further includes:

[0076] Use the cosine similarity algorithm to calculate the similarity between the target voiceprint feature vector and the voiceprint template in the database.

[0077] If the similarity is greater than the threshold, the target voiceprint feature vector passes the verification; if the similarity is lower than the threshold, it is considered that there may be a problem with the target voiceprint feature vector, and further processing or re - collection of voice data is performed to ensure the accuracy of feature extraction.

[0078] Specifically, the first voiceprint feature extraction includes the processing of the first few layers of a deep neural network (such as TDNN) to obtain a voiceprint representation at a higher level than the audio signal features; the second voiceprint feature extraction includes the subsequent layer processing of the deep neural network (such as the x - vector system), PLDA (Probabilistic Linear Discriminant Analysis) processing, etc., to obtain a target voiceprint feature vector with a fixed dimension, which is suitable for voiceprint comparison and recognition.

[0079] In the embodiments of the present invention, the extracting audio signal features from the target voice data includes:

[0080] Performing frame - splitting processing on the target voice data to obtain audio signal data;

[0081] Performing pre - emphasis and adding a Hamming window to each frame of the audio signal data to obtain target audio signal data;

[0082] Performing energy extraction on the target audio signal data to obtain frame - based energy eigenvalue;

[0083] Performing filtering processing on the frame - based energy eigenvalue to obtain frame - based audio signal features, and aggregating all the frame - based audio signal features to obtain the audio signal features.

[0084] Specifically, continuous voice data signals are segmented into multiple short time periods (frames), usually with a frame length of 20 - 30 ms and a frame shift of 10 ms, to obtain a series of short - time audio signal data segments (frames). In order to enhance high - frequency components and balance the speech spectrum, a first - order high - pass filter is applied to multiply each frame of the signal by a Hamming window function, and the energy of the audio signal data after adding the Hamming window is obtained, that is, the energy value of each frame of the signal is calculated. The frame - based energy eigenvalue can reflect the intensity information of the speech.

[0085] The present invention usually uses a Mel filter bank or a Bark filter bank to process the frame - based energy eigenvalue, obtain the spectrum of each frame of audio signal data, and take the logarithmic energy output by a group of triangular filters as the frequency - divided audio signal features.

[0086] Exemplarily, in the medical and health scenario, a deep - learning - driven speech recognition model and natural language processing technology can be adopted for health insurance recommendation and purchase, while supporting multi - turn conversations and intent recognition.

[0087] Among them, the user can input health needs through voice, such as "I am 30 years old, often stay up late, and need a physical examination package". The system parses the voice content, can frame it with a frame length of 25 ms and a frame shift of 10 ms, and performs pre-emphasis and Hamming window addition processing on each frame of audio signal data using a pre-emphasis filter with α = 0.97 and a window function. Then, calculate the energy value of each frame, process it using a 40-channel Mel filter bank, output the audio signal features and key information such as age, living habits, and health goals, and match the medical guidelines according to the user's age, gender, family medical history, living habits, etc. to dynamically generate a personalized physical examination item list. Then, combined with the insurance product database, recommend an insurance plan that covers the user's needs, such as "a combination product including critical illness insurance + medical insurance".

[0088] Specifically, based on collaborative filtering, content recommendation or reinforcement learning algorithms, combined with the user's historical behavior (such as past physical examination records, insurance purchase preferences), calculate the consumption intention matrix in real time, dynamically adjust the recommendation list, and automatically generate an order after the user's voice confirmation.

[0089] In the embodiment of the present invention, the voiceprint features are gradually refined through multi-level processing, from the original signal to the audio features, then to the preliminary voiceprint representation, and finally a highly discriminative voiceprint feature vector is obtained, improving the accuracy and robustness of voiceprint recognition.

[0090] S3. Perform phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector.

[0091] In the embodiment of the present invention, the phoneme feature extraction of the target voice data is realized by performing unit division, phoneme recognition, phoneme decoding, and smoothing filtering on the target voice data.

[0092] In the embodiment of the present invention, the performing phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector includes:

[0093] Perform unit division on the target voice data to obtain language phoneme units;

[0094] Perform phoneme recognition on the language phoneme units to obtain corresponding phoneme probability feature values;

[0095] Perform phoneme combination decoding according to the phoneme probability feature values to obtain a phoneme original vector;

[0096] Perform smoothing filtering on the phoneme original vector to obtain a target phoneme decoding vector.

[0097] Specifically, voice activity detection can be used to remove the silent segments and non-speech noises of the target voice data, and an automatic phoneme segmentation technology is used to segment the continuous target voice data into speech phoneme units according to a preset phoneme dictionary.

[0098] The present invention uses a deep neural network (DNN), a convolutional neural network (CNN), or a Transformer model (such as Wav2Vec 2.0) for phoneme recognition, calculates the probability distribution of each phoneme unit, and the phoneme probability eigenvalue reflects its confidence in the speech data, indicating the possibility that the frame belongs to different phonemes.

[0099] Specifically, connectionist temporal classification (CTC) or a recurrent neural network (RNN) decoder can be adopted, and combined with the Viterbi algorithm such as the Viterbi Algorithm to convert the discrete phoneme probability eigenvalues into continuous phoneme raw vectors.

[0100] The smoothing filtering in the present invention refers to eliminating noise and mutations in the phoneme raw vectors, improving the robustness of the vectors. The smoothing filtering can adopt a moving average filter (Moving Average Filter) to smooth adjacent phoneme vectors, and can also remove outliers based on statistics, such as removing phonemes with extremely low probabilities, so as to obtain the final target phoneme decoding vector.

[0101] Exemplarily, in the scenario of purchasing drugs and medical devices, a user may interact with intelligent terminals such as smart speakers and mobile phone voice assistants through voice to query information about drugs or medical devices, perform purchase operations, etc. In order to better understand the user's voice commands, it is necessary to extract phoneme features from the user's voice data, and then realize speech recognition and intention understanding.

[0102] Specifically, when the user speaks a voice command related to the purchase of drugs and medical devices in front of the intelligent terminal, the voice acquisition module of the intelligent terminal will collect the user's voice signal in real time and convert it into a digital signal to obtain the target voice data. For example, when the user says: "I want to buy a box of ibuprofen sustained-release capsules." Using voice signal processing technologies such as short-time energy analysis and zero-crossing rate analysis, the target voice data is segmented, and the continuous voice signal is divided into relatively independent voice units. For example, the above voice "I want to buy a box of ibuprofen sustained-release capsules" is divided into multiple phoneme units, such as phoneme segments corresponding to "wǒ (I)", "xiǎng (want)", "gòu (purchase)", "mǎi (buy)", etc.

[0103] Specifically, a pre-trained phoneme recognition model such as a deep learning-based acoustic model is used to recognize each language phoneme unit. The model will output the probability distribution of the phoneme unit belonging to each possible phoneme according to the input phoneme unit features, that is, the phoneme probability eigenvalue. The phoneme topology decoding algorithm (such as the Viterbi algorithm) is adopted, and combined with a language model such as a statistical language model based on the vocabulary in the field of drugs and medical devices, phoneme combination decoding is performed according to the phoneme probability eigenvalue.

[0104] The target phoneme decoding vector is input into the speech recognition system and further converted into the corresponding text information, i.e., "I want to buy a box of ibuprofen sustained-release capsules". Semantic analysis is performed on the recognized text information to understand the user's purchase intention and determine information such as the name and quantity of the drug the user wants to buy. According to the user's purchase intention, relevant information such as the price, inventory, and manufacturer of the drug is queried in the drug and medical device database, and the query results are fed back to the user or the user is guided to complete the purchase process.

[0105] Through the above implementation steps, in the drug and medical device purchase scenario, the phoneme features of the user's voice commands can be accurately extracted and speech recognition can be performed, thus realizing intelligent voice interaction and business processing.

[0106] In the embodiment of the present invention, the phoneme recognition error rate is reduced through four-level progressive processing, and the decoding accuracy is greatly improved. A hierarchical architecture is used to perform phoneme decoding on the target voice data to improve the decoding efficiency.

[0107] S4. Feature fusion is performed according to the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector.

[0108] In the embodiment of the present invention, by aligning the dimensions of the target voiceprint feature vector and the target phoneme decoding vector, and then splicing the aligned feature vectors, a spliced speech feature is obtained. The spliced speech feature is mapped to a unified semantic space through a non-linear transformation to obtain a comprehensive semantic vector.

[0109] In the embodiment of the present invention, the obtaining of the comprehensive semantic vector by performing feature fusion according to the target voiceprint feature vector and the target phoneme decoding vector includes:

[0110] Perform vector dimension alignment on the target voiceprint feature vector and the target phoneme decoding vector to obtain an aligned voiceprint feature and an aligned phoneme feature;

[0111] Perform feature splicing on the aligned voiceprint feature and the aligned phoneme feature to obtain a spliced speech feature;

[0112] Map the spliced speech feature to a unified semantic space to obtain a comprehensive semantic vector.

[0113] Specifically, the vector dimension alignment can solve the problem of inconsistent dimensions in different feature spaces. The target voiceprint feature vector and the target phoneme decoding vector can be dimensionally reduced through a fully connected layer, and causal convolution (Causal Conv) can also be used for temporal compression to achieve vector dimension alignment.

[0114] The present invention can use a Transformer encoder to perform a non-linear transformation on the spliced speech features to map them to a unified semantic space.

[0115] In an embodiment of the present invention, through a dynamic dimension alignment algorithm, lossless fusion of voiceprint and phoneme features is achieved, improving the fusion ability of cross-feature information, and at the same time improving the accuracy and precision of speech data recognition.

[0116] S5. Construct a consumption intention matrix of the target user according to the comprehensive semantic vector, and make a consumption decision according to the consumption intention matrix to obtain an optimal consumption plan.

[0117] In an embodiment of the present invention, the consumption intention matrix is a two-dimensional array, which is used to represent the intention intensity of the target user for different consumer goods, and can intuitively display the consumption willingness of the target user for different consumer goods, providing a basis for subsequent consumption decisions.

[0118] Exemplarily, in a fintech business scenario, the consumption intention matrix is used to accurately match the wealth management needs of the target user with the characteristics of financial products, achieving a balance between maximizing returns and controlling risks. The set of consumer goods can include financial instruments such as deposit products, funds, insurance, and credit. The text description embedding uses a RoBERTa model trained with financial professional corpus. For example, a "hybrid fund" is parsed into feature vectors such as "equity-debt balance" and "medium risk"; a "mortgage business loan" is associated with dimensions such as "LPR + 50BP" and "repayable at any time".

[0119] Among them, the comprehensive semantic vector of the target user is constructed by integrating the balance sheet (such as property valuation, debt ratio), risk assessment results (such as R3 balanced type), and behavioral data (such as frequently checking the gold market). When calculating feature similarity, compliance constraints are introduced. For example, for conservative users, stock funds are automatically filtered (the similarity threshold is set to only display products with a maximum drawdown < 5%), while for high-net-worth customers, the weight of private equity products is increased.

[0120] Specifically, in the generated consumption intention matrix, the row dimension divides financial goals such as "asset appreciation", "risk hedging", and "liquidity management", the column dimension is associated with specific financial products, and the matrix values comprehensively consider quantitative indicators such as historical annualized return, Sharpe ratio, and subscription and redemption costs.

[0121] In the consumption decision-making stage, based on the consumption intention matrix, a product portfolio that fits the user's financial plan is selected, and preferential strategies are intelligently superimposed, which may include new customer exclusive rights (such as a 0.1% discount on the subscription fee), asset scale ladder rewards (such as exemption from account management fees for AUM exceeding 1 million), cross-product collaborative discounts (such as doubling credit card points for purchasing life insurance), etc. An optimal portfolio return can be pursued under a given risk exposure through a multi-objective optimization model.

[0122] As shown Figure 4 In the embodiment of the present invention, constructing the consumption intention matrix of the target user according to the comprehensive semantic vector includes:

[0123] Obtain a set of consumption items, perform text description embedding on the set of consumption items, and obtain corresponding item semantic features;

[0124] Calculate the feature similarity between the item semantic features and the comprehensive semantic vector;

[0125] Generate the consumption intention matrix of the target user according to the consumption items whose feature similarity meets the preset threshold.

[0126] When calculating the feature similarity in the present invention, historical shopping data and browsing behaviors of users can also be introduced, combined with association rule analysis, to more accurately predict the consumption intention of users, and the consumption intention matrix can also be adjusted in real time according to the real-time behaviors and feedback of users.

[0127] Therefore, after step S5, it further includes:

[0128] Send an invitation for feedback on consumption decisions to the user side;

[0129] Collect the feedback on the consumption decision results sent by the user side, and optimize the consumption decision model according to the feedback.

[0130] Specifically, the text description embedding is a technology that converts text data into a numerical vector form for computer processing and analysis. A pre-trained language model (such as BERT, Sentence-BERT) can be used to encode texts such as the titles, descriptions, and attribute labels of the set of consumption items to generate multi-dimensional semantic vectors, thereby generating an item semantic feature matrix. The item semantic features can reflect the semantic relationships and similarities between items.

[0131] The present invention can calculate the feature similarity between the item semantic features and the comprehensive semantic vector through cosine similarity. The feature similarity can measure the matching degree between the item and the consumption intention of the user, and then select high-similarity consumption items through a preset similarity threshold.

[0132] Specifically, generate a consumption intention matrix according to the item IDs and similarity values of the high-similarity consumption items. The rows of the consumption intention matrix can be the comprehensive semantic vector and the corresponding similarity, and the columns of the consumption intention matrix can be the high-similarity consumption items and the corresponding similarity. The consumption intention matrix can intuitively display the consumption willingness of the target user for different items and provide a basis for subsequent consumption decisions.

[0133] In the embodiments of the present invention, making a consumption decision according to the consumption intention matrix to obtain an optimal consumption plan includes:

[0134] Selecting the consumer goods of the target user according to the consumption intention matrix;

[0135] Obtaining the available preferential information of the target user and the consumer goods, and generating an optimal consumption plan for the consumer goods according to the preferential information.

[0136] Specifically, vertical selection and horizontal filtering can be performed according to the consumption intention matrix. The vertical selection refers to selecting several consumer goods with the highest similarity in the intention dimension, and the horizontal filtering refers to excluding the goods on the blacklist of the target user or the goods with low inventory.

[0137] Among them, the available preferential information includes the consumer vouchers of the target user and the vouchers of the consumer goods. The branch and bound method can be used to optimize the combination of the available preferential information, so as to generate an optimal consumption plan, help the target user maximize the consumption utility, and improve the consumption utility.

[0138] Exemplarily, in the medical and health scenario, the construction and decision-making process of the consumption intention matrix focuses on the personalized diagnosis and treatment needs of patients and the optimization of medical costs. First, the consumer goods can include drugs, medical devices, examination items and health services (such as online consultations). The BERT model enhanced by medical knowledge is used to embed the item descriptions: for example, the drug name "Aspirin Enteric-coated Tablets" is encoded into a vector containing semantics such as "antiplatelet aggregation" and "cardiovascular disease"; the examination item "Coronary CT" is associated with features such as "coronary heart disease screening" and "non-invasive imaging".

[0139] Among them, the comprehensive semantic vector of the patient is generated by fusing his electronic health record (such as past medical history, allergy record), recent symptom description (such as "chest pain for 3 days") and insurance coverage (such as medical insurance type). When calculating the feature similarity, items with high medical necessity are preferentially matched. For example, for patients with angina pectoris, the similarity weight of nitroglycerin tablets is significantly higher than that of health care products, and drugs outside the medical insurance catalog are dynamically filtered (the threshold is set to cover 80% of the basic drugs).

[0140] Specifically, in the generated consumption intention matrix, the row dimension includes medical intentions such as "acute treatment", "long-term medication", and "preventive examination", the column dimension is associated with specific drugs and services, and the matrix value reflects the clinical matching degree and cost-benefit score.

[0141] In the consumer decision-making stage, drugs combinations strongly associated with the intention matrix (such as "aspirin + statins") are preferentially recommended, and multi-source preferential information is accessed in real time, such as the overall medical insurance payment ratio (e.g., 60% reimbursement in tertiary hospitals), pharmaceutical company patient assistance programs (e.g., buy 3 boxes and get 1 box free), and discount packages for inspection items (e.g., a 30% discount on the combined price of electrocardiogram + myocardial enzyme tests). Through a constrained optimization algorithm, an optimal payment path is automatically generated on the premise of meeting the effectiveness of the diagnosis and treatment plan. For example, for diabetic patients, the combined plan may be: replacing the original drug with a generic drug that has passed the consistency evaluation (saving 40%), using the special outpatient disease quota of medical insurance to pay for blood glucose test strips, and simultaneously activating the annual discount for the contracted family doctor service.

[0142] In the embodiments of the present invention, by constructing a consumption intention matrix, the consumption preferences and needs of target users can be captured and analyzed more accurately, so as to recommend products or services that better meet their personalized needs to users, improving the consumption decision-making efficiency and enhancing the accuracy of consumption decisions.

[0143] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0144] As Figure 5 shown, it is a functional module diagram of an intelligent shopping decision-making device for speech recognition provided by an embodiment of the present invention.

[0145] In the embodiments of the present disclosure, an intelligent shopping decision-making device for speech recognition is provided. The intelligent shopping decision-making device for speech recognition corresponds one-to-one with the above-mentioned intelligent shopping decision-making method based on speech recognition. As Figure 5 shown, the intelligent shopping decision-making device 100 for speech recognition can be installed in an electronic device. According to the functions implemented, the intelligent shopping decision-making device 100 for speech recognition includes a speech enhancement module 101, a feature extraction module 102, a phoneme decoding module 103, a feature fusion module 104, and a consumption decision module 105. The detailed descriptions of each functional module are as follows:

[0146] The speech enhancement module 101 is used to obtain the voice shopping data of the target user, perform speech enhancement on the voice shopping data, and obtain target voice data;

[0147] The feature extraction module 102 is used to extract the voiceprint feature of the target voice data to obtain a target voiceprint feature vector;

[0148] The phoneme decoding module 103 is used to perform phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector;

[0149] The feature fusion module 104 is used to perform feature fusion based on the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector;

[0150] The consumption decision module 105 is used to construct a consumption intention matrix of the target user according to the comprehensive semantic vector, and perform a consumption decision according to the consumption intention matrix to obtain an optimal consumption plan.

[0151] In one embodiment, when the voice enhancement module 101 performs voice enhancement on the voice shopping data to obtain target voice data, it is used for:

[0152] Perform signal decomposition on the voice shopping data to obtain a noise spectrum feature and a user voice spectrum feature;

[0153] Construct an adaptive filter according to the noise spectrum feature and the parameters of the pre-acquired adaptive filter;

[0154] Perform noise suppression on the user voice spectrum feature according to the adaptive filter to obtain target voice data.

[0155] In one embodiment, when the voice enhancement module 101 performs signal decomposition on the voice shopping data to obtain a noise spectrum feature and a user voice spectrum feature, it is used for:

[0156] Perform signal conversion on the voice shopping data according to the short-time Fourier transform to obtain a time-frequency domain signal;

[0157] Extract a noise spectrum basis matrix and a voice spectrum basis matrix in the voice shopping data according to the time-frequency domain signal;

[0158] Perform feature extraction on the noise spectrum basis matrix and the voice spectrum basis matrix respectively to obtain a noise spectrum feature and a user voice spectrum feature.

[0159] In one embodiment, when the feature extraction module 102 performs voiceprint feature extraction on the target voice data to obtain a target voiceprint feature vector, it is used for:

[0160] Perform audio feature extraction on the target voice data to obtain an audio signal feature;

[0161] Perform first voiceprint feature extraction on the audio signal feature to obtain a preliminary voiceprint feature;

[0162] Perform second voiceprint feature extraction on the preliminary voiceprint feature to obtain a target voiceprint feature vector.

[0163] In one embodiment, when the feature extraction module 102 performs audio feature extraction on the target voice data to obtain an audio signal feature, it is used for:

[0164] Perform frame segmentation on the target voice data to obtain audio signal data;

[0165] Perform pre-emphasis and add a Hamming window to each frame of the audio signal data to obtain target audio signal data;

[0166] Perform energy extraction on the target audio signal data to obtain frame energy eigenvalues;

[0167] Perform filtering on the frame energy eigenvalues to obtain frame audio signal features, and aggregate all the frame audio signal features to obtain the audio signal features.

[0168] In one embodiment, when the phoneme decoding module 103 performs phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector, it is used for:

[0169] Perform unit division on the target voice data to obtain language phoneme units;

[0170] Perform phoneme recognition on the language phoneme units to obtain corresponding phoneme probability eigenvalues;

[0171] Perform phoneme combination decoding according to the phoneme probability eigenvalues to obtain a phoneme original vector;

[0172] Perform smoothing filtering on the phoneme original vector to obtain a target phoneme decoding vector.

[0173] In one embodiment, when the feature fusion module 104 performs feature fusion according to the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector, it is used for:

[0174] Perform vector dimension alignment on the target voiceprint feature vector and the target phoneme decoding vector to obtain aligned voiceprint features and aligned phoneme features;

[0175] Perform feature splicing on the aligned voiceprint features and the aligned phoneme features to obtain spliced speech features;

[0176] Map the spliced speech features to a unified semantic space to obtain a comprehensive semantic vector.

[0177] In one embodiment, when the consumption decision module 105 constructs the consumption intention matrix of the target user according to the comprehensive semantic vector, it is used for:

[0178] Obtain a set of consumer goods, perform text description embedding on the set of consumer goods to obtain corresponding item semantic features;

[0179] Calculate the feature similarity between the semantic features of the item and the comprehensive semantic vector;

[0180] Generate the consumption intention matrix of the target user based on the consumer goods whose feature similarity meets a preset threshold.

[0181] In one embodiment, when the consumption decision-making module 105 performs consumption decision-making according to the consumption intention matrix to obtain an optimal consumption plan, it is used for:

[0182] Select the consumer goods of the target user according to the consumption intention matrix;

[0183] Obtain the available preferential information of the target user and the consumer goods, and generate an optimal consumption plan for the consumer goods according to the preferential information.

[0184] In the present invention, the specific limitations of an intelligent shopping decision-making device for speech recognition can refer to the limitations of an intelligent shopping decision-making method based on speech recognition in the above text, which will not be elaborated here. Each module in the above intelligent shopping decision-making device for speech recognition can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so as to facilitate the processor to call and execute the operations corresponding to the above respective modules.

[0185] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 6 shown. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of an intelligent shopping decision-making method based on speech recognition.

[0186] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 7As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of an intelligent shopping decision-making method based on speech recognition.

[0187] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:

[0188] Obtain the voice shopping data of the target user, perform voice enhancement on the voice shopping data to obtain target voice data;

[0189] Extract the voiceprint feature of the target voice data to obtain a target voiceprint feature vector;

[0190] Perform phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector;

[0191] Perform feature fusion according to the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector;

[0192] Construct a consumption intention matrix of the target user according to the comprehensive semantic vector, and perform consumption decision-making according to the consumption intention matrix to obtain an optimal consumption plan.

[0193] In several embodiments provided by the present invention, it should be understood that the disclosed devices and apparatuses can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0194] In addition, each functional module in each embodiment of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a hardware plus a software functional module.

[0195] Therefore, in all respects, the embodiments should be considered exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Thus, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced by the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0196] For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and the present invention can be implemented in other specific forms without departing from the spirit or basic characteristics of the present invention.

[0197] In some embodiments of this embodiment, a computer-readable storage medium is provided, on which a computer program is stored. It is characterized in that when the computer program is executed by a processor, the steps of the method described in the above embodiment are implemented.

[0198] The readable storage medium of the present invention stores a computer program, and when the computer program is executed by a processor of an electronic device, it can implement:

[0199] Obtain the voice shopping data of the target user, perform voice enhancement on the voice shopping data to obtain target voice data;

[0200] Extract the voiceprint feature of the target voice data to obtain a target voiceprint feature vector;

[0201] Perform phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector;

[0202] Perform feature fusion according to the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector;

[0203] Construct a consumption intention matrix of the target user according to the comprehensive semantic vector, and make a consumption decision according to the consumption intention matrix to obtain an optimal consumption plan.

[0204] It should be noted that for the functions or steps that can be realized by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0205] The computer-readable storage medium may also store at least one computer-executable program / instructions, such as computer-readable instructions. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory, etc. The computer-readable storage medium may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, the non-transitory computer-readable storage medium may be connected to a computing device such as a computer. Then, when the computing device runs the computer-readable instructions stored on the computer-readable storage medium, the various methods described above may be performed.

[0206] In addition, the computer device may also include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (such as a keyboard, a mouse, a speaker, etc.).

[0207] The processor may communicate with external devices via the I / O bus through a wired or wireless network.

[0208] In one embodiment, the at least one computer-executable instruction may also be compiled into or constitute a software product / computer program product, and when one or more computer-executable instructions are run by a processor, the steps of the various functions and / or methods in the embodiments described in the present technology are performed.

[0209] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0210] Those skilled in the art can clearly understand that for the convenience and brevity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0211] In the embodiments provided in the present disclosure, it should be understood that the disclosed device and method can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the block can occur in a different order from that marked in the accompanying drawings. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0212] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.

[0213] It should be noted that if non-company software tools or components appear in the embodiments of this application, they are only used for illustrative introduction and do not represent actual use.

Claims

1. An intelligent shopping decision-making method based on speech recognition, characterized in that The method includes: Obtain the voice shopping data of the target user, perform voice enhancement on the voice shopping data to obtain target voice data; Extract the voiceprint feature from the target voice data to obtain a target voiceprint feature vector; Perform phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector; Perform feature fusion according to the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector; Construct a consumption intention matrix of the target user according to the comprehensive semantic vector, and make a consumption decision according to the consumption intention matrix to obtain an optimal consumption plan.

2. The intelligent shopping decision-making method based on speech recognition according to claim 1, characterized in that The performing voice enhancement on the voice shopping data to obtain target voice data includes: Perform signal decomposition on the voice shopping data to obtain a noise spectrum feature and a user voice spectrum feature; Construct an adaptive filter according to the noise spectrum feature and the parameters of the pre-acquired adaptive filter; Suppress noise on the user voice spectrum feature according to the adaptive filter to obtain target voice data.

3. The intelligent shopping decision-making method based on speech recognition according to claim 2, characterized in that, The performing signal decomposition on the voice shopping data to obtain a noise spectrum feature and a user voice spectrum feature includes: Perform signal conversion on the voice shopping data according to the short-time Fourier transform to obtain a time-frequency domain signal; Extract a noise spectrum basis matrix and a voice spectrum basis matrix in the voice shopping data according to the time-frequency domain signal; Perform feature extraction on the noise spectrum basis matrix and the voice spectrum basis matrix respectively to obtain a noise spectrum feature and a user voice spectrum feature.

4. The intelligent shopping decision-making method based on speech recognition according to claim 1, wherein The extracting the voiceprint feature from the target voice data to obtain a target voiceprint feature vector includes: Extract audio feature from the target voice data to obtain an audio signal feature; Perform first voiceprint feature extraction on the audio signal feature to obtain a preliminary voiceprint feature; Perform second voiceprint feature extraction on the preliminary voiceprint feature to obtain a target voiceprint feature vector.

5. The intelligent shopping decision-making method based on speech recognition according to claim 1, wherein The constructing the consumption intention matrix of the target user according to the comprehensive semantic vector includes: Obtain a set of consumption items, perform text description embedding on the set of consumption items to obtain corresponding item semantic features; Calculate the feature similarity between the item semantic feature and the comprehensive semantic vector; Generate a consumption intention matrix of the target user according to the consumption items whose feature similarity meets a preset threshold.

6. The intelligent shopping decision-making method based on speech recognition according to claim 1, characterized in that The performing phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector includes: Perform unit division on the target voice data to obtain language phoneme units; Perform phoneme recognition on the language phoneme units to obtain corresponding phoneme probability feature values; Perform phoneme combination decoding according to the phoneme probability feature values to obtain a phoneme original vector; Perform smoothing filtering on the phoneme original vector to obtain a target phoneme decoding vector.

7. The intelligent shopping decision-making method based on speech recognition according to claim 1, characterized in that, The performing feature fusion according to the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector includes: Perform vector dimension alignment on the target voiceprint feature vector and the target phoneme decoding vector to obtain an aligned voiceprint feature and an aligned phoneme feature; Perform feature splicing on the aligned voiceprint features and the aligned phoneme features to obtain spliced speech features; Map the spliced speech features to a unified semantic space to obtain a comprehensive semantic vector.

8. An intelligent shopping decision-making device for speech recognition, characterized in that, The device includes: A voice enhancement module, configured to obtain voice shopping data of a target user, perform voice enhancement on the voice shopping data to obtain target voice data; A feature extraction module, configured to extract voiceprint features from the target voice data to obtain a target voiceprint feature vector; A phoneme decoding module, configured to perform phoneme topology decoding on the target voice data to obtain a target phoneme decoding vector; A feature fusion module, configured to perform feature fusion based on the target voiceprint feature vector and the target phoneme decoding vector to obtain a comprehensive semantic vector; A consumption decision module, configured to construct a consumption intention matrix of the target user according to the comprehensive semantic vector, and perform a consumption decision according to the consumption intention matrix to obtain an optimal consumption plan.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor, so that the at least one processor can execute a voice recognition-based intelligent shopping decision method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements a voice recognition-based intelligent shopping decision method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Agricultural product wholesale order automatic generation method and device, equipment and medium

    CN121544346A