Intelligent voice-based fault diagnosis method, device and equipment and readable storage medium
By combining intelligent voice and vision technologies, a fault diagnosis method has been developed that solves the problem of elderly people having difficulty diagnosing faults in household appliances. It provides convenient and efficient fault diagnosis services and is applicable to home management devices and smart terminals.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHONGSHAN POWER SUPPLY BUREAU OF GUANGDONG POWER GRID
- Filing Date
- 2026-05-27
- Publication Date
- 2026-07-14
AI Technical Summary
Elderly people find it difficult to diagnose household appliance malfunctions by searching for text information and operating electronic devices, and existing technologies cannot provide convenient fault diagnosis methods.
A fault diagnosis method based on intelligent voice is adopted. By acquiring the user's voice query signal, combining vision technology and fault diagnosis model, fault diagnosis results are generated and optimized into voice interaction signals to provide convenient fault diagnosis services.
It improves the efficiency and accuracy of fault diagnosis, lowers the operational threshold, is especially suitable for the elderly, reduces reliance on professional maintenance personnel, and saves time and economic costs.
Smart Images

Figure CN122392530A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of fault diagnosis technology, and more specifically, to a fault diagnosis method, apparatus, device, and readable storage medium based on intelligent voice. Background Technology
[0002] In modern life, home appliances have become deeply integrated into many households, becoming an indispensable part of daily life. However, due to their high frequency of use, many home appliances inevitably break down. For some relatively simple damages, family members with some knowledge can repair them themselves by consulting relevant information, without needing to hire a repairman.
[0003] However, it should be noted that elderly people face significant difficulties in diagnosing and troubleshooting problems using text-based information. Due to factors such as declining eyesight, unfamiliarity with electronic devices, or limited literacy, older adults may find it challenging to effectively utilize existing text-based search methods to diagnose household appliance malfunctions.
[0004] Therefore, how to develop a method specifically for the elderly to help them easily diagnose malfunctions of household appliances has become a key concern for those skilled in the art. Summary of the Invention
[0005] In view of this, this application provides a fault diagnosis method, apparatus, device and readable storage medium based on intelligent voice, for helping the elderly diagnose faults in household appliances.
[0006] To achieve the above objectives, the following solution is proposed: A fault diagnosis method based on intelligent voice includes: The system acquires user-input voice query signals and a trained fault diagnosis model, wherein the voice query signals include the type of faulty product and the form of fault manifestation. By combining visual technology, the visual information of the product corresponding to the type of faulty product is determined; The voice query signal is converted into text to generate text fault information; The fault diagnosis model is used to perform fault diagnosis based on the text fault information and the visual information, and to generate fault diagnosis results and repair solutions. The fault diagnosis results and the repair plan are optimized using natural language processing technology, and text content is generated. Determine the voice features corresponding to the user; Based on the speech features, the text content is converted into a speech interaction signal, and the speech interaction signal is played.
[0007] Optionally, acquiring the user-input voice query signal includes: Respond to the user's voice input and perform intent recognition on the input voice; When it is determined that the input voice corresponds to a fault diagnosis, all data types contained in the input voice are determined; Each data type is matched with a preset fault diagnosis template to determine supplementary data types. The fault diagnosis template is used to characterize all types of identifiers involved in fault diagnosis. Based on the aforementioned supplementary data types, an interactive voice signal is generated; Play the interactive voice signal to obtain the user's feedback voice signal; The feedback voice signal and the input voice signal are integrated to obtain the voice query signal.
[0008] Optionally, integrating the feedback voice signal and the input voice to obtain the voice query signal includes: Obtain a filter adjusted by the least mean square error or least squares algorithm, and use the filter to denoise the feedback speech signal and the input speech; By combining a speech enhancement model with an attention mechanism and a multi-scale convolutional speech network, speech enhancement is performed on the denoised feedback speech signal and the input speech. The voice query signal is obtained by splicing the enhanced feedback voice signal and the input voice.
[0009] Optionally, obtain the trained fault diagnosis model, including: Acquire training samples of different types of household appliances. Each training sample contains the corresponding household appliance's training diagnosis type, training fault manifestation form, training visual data, and training solution. Construct an initial diagnostic model; The initial diagnostic model is trained using each training sample, and the final initial diagnostic model is used as the trained fault diagnosis model.
[0010] Optionally, obtaining training samples of different types of household appliances includes: Combine smart cameras to collect training visual data for different types of home appliances and different training and diagnostic types; By using web crawling technology, we can obtain the training fault manifestations and training solutions for different types of household appliances and different training diagnostic types. For different household appliances, each training diagnostic type of the household appliance is integrated with the corresponding training visual data, training fault manifestations, and training solutions to form a training sample.
[0011] Optionally, training the initial diagnostic model using each training sample includes: Each training sample is sequentially input into the initial diagnostic model to obtain the prediction scheme and prediction diagnosis result corresponding to each training sample; Based on each training sample and its corresponding prediction scheme and prediction diagnosis result, the loss value of the initial diagnosis model is calculated; Based on the loss value, the parameters of the initial diagnostic model are adjusted using the gradient descent method. Based on the loss value of each training sample, set the training frequency for each training sample; The initial diagnostic model, after parameter adjustment, is trained based on each training sample and its corresponding training frequency.
[0012] Optionally, the step of using natural language processing technology to optimize the fault diagnosis results and the repair plan, and generating text content, includes: The fault diagnosis results and the repair plan are segmented into characters to obtain multiple fields; Fields that are non-linguistic components are removed, and fields that are abbreviations are expanded to obtain multiple optimized fields; The various optimized fields are integrated to generate text content.
[0013] A fault diagnosis device based on intelligent voice, comprising: The acquisition module is used to acquire the user-input voice query signal and the trained fault diagnosis model. The voice query signal includes the faulty product type and the fault manifestation. The module is used to combine visual technology to determine the visual information of the product corresponding to the faulty product type. The conversion module is used to convert the voice query signal into text and generate text fault information; The generation module is used to perform fault diagnosis based on the text fault information and the visual information using the fault diagnosis model, and generate fault diagnosis results and repair solutions. An optimization module is used to optimize the fault diagnosis results and the repair plan using natural language processing technology, and to generate text content. A determination module is used to determine the voice features corresponding to the user; The playback module is used to convert the text content into a voice interaction signal based on the voice features, and to play the voice interaction signal.
[0014] A fault diagnosis device based on intelligent voice, including a memory and a processor; The memory is used to store programs; The processor is used to execute the program to implement the various steps of the above-described intelligent voice-based fault diagnosis method.
[0015] A readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the various steps of the above-described intelligent voice-based fault diagnosis method.
[0016] As can be seen from the above technical solution, the fault diagnosis method based on intelligent voice provided in this application can acquire the user-inputted voice query signal and a trained fault diagnosis model. The voice query signal includes the type of faulty product and the manifestation of the fault. Therefore, in the process of using this application, only the user needs to input the voice query signal, eliminating the need for manual text input. This is especially convenient for people who are not proficient in text input or operating electronic devices, such as the elderly. For example, the elderly can initiate a query simply by stating the type of faulty household appliance and its manifestation without the effort of typing, saving time and simplifying the operation, thus greatly improving the efficiency of fault diagnosis. Simultaneously, this application can combine visual technology to determine the visual information of the product corresponding to the faulty product type; perform text conversion on the voice query signal to generate text fault information; and use the fault diagnosis model to perform fault diagnosis based on the text fault information and the visual information to generate fault diagnosis results and repair solutions. Based on this, this application can integrate multimodal information for fault diagnosis, improving the accuracy of fault diagnosis and repair solution determination. On the one hand, it utilizes the text fault information containing the faulty product type and fault manifestation after voice conversion; on the other hand, it combines the product visual information obtained through visual technology to provide data support for the fault diagnosis model from different dimensions, allowing the fault diagnosis model to analyze the fault more comprehensively and accurately, thereby matching the corresponding repair solutions. Subsequently, this application can use natural language processing technology to optimize the fault diagnosis results and the repair plan, and generate text content; determine the user's corresponding voice characteristics; convert the text content into a voice interaction signal based on the voice characteristics, and play the voice interaction signal; based on this, this application can use natural language processing technology to optimize the fault diagnosis results and repair plan, transforming professional and complex diagnosis and repair content into easy-to-understand text content, making it easier for users to understand; at the same time, the voice interaction signal played by this application is generated in combination with the user's corresponding voice characteristics, conforming to the user's voice habits, enhancing the interactive experience between the user and the system, especially for specific user groups such as the elderly, voice output is easier to accept and understand than text reading. Therefore, this application can integrate voice interaction, visual technology assistance, natural language optimization, and personalized voice output functions to provide fault diagnosis and repair plan determination services for the elderly who are not good at text search and operating electronic devices, as well as ordinary users who have difficulty understanding professional terminology, thus expanding the applicable population of fault diagnosis services. At the same time, it lowers the threshold for ordinary users to troubleshoot faults on their own. Common faults can be initially located and handled without the need for professional repair personnel to come to the site. It can help users quickly solve simple faults, saving time and economic costs for on-site repairs. It can also help professional repair personnel to make fault predictions in advance, prepare the corresponding tools and parts in advance, and improve the efficiency of on-site repairs. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0018] Figure 1 This is a flowchart of a fault diagnosis method based on intelligent voice disclosed in an embodiment of this application; Figure 2 This is a structural block diagram of a fault diagnosis device based on intelligent voice disclosed in an embodiment of this application; Figure 3 This is a hardware structure block diagram of a fault diagnosis device based on intelligent voice disclosed in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] This application provides a fault diagnosis method based on intelligent voice. This method can be applied to various home management devices or intelligent voice systems, as well as to various computer terminals or smart terminals. The executing entity can be the processor or server of the computer terminal or smart terminal.
[0021] Various home concierge or smart voice systems may include microphones, speakers, and smart cameras. Microphones and speakers are used to capture user voice input and play system feedback or guidance voice.
[0022] Next, combine Figure 1 The fault diagnosis method based on intelligent voice of this application is described in detail, including the following steps: Step S1: Obtain the user's voice query signal and the trained fault diagnosis model.
[0023] Specifically, the voice query signal may include the type of faulty product and the form of the fault.
[0024] The faulty product type can be an appliance, such as a rice cooker, water dispenser, microwave oven, or electric steamer.
[0025] The faults may manifest as abnormal heating, loud operating noise, display malfunction, overflow, abnormal vibration, and leakage.
[0026] Different types of electrical appliances can exhibit different fault manifestations.
[0027] For example, the malfunctions of a rice cooker can manifest as overflowing, not cooking rice properly, and leakage of electricity. For example, microwave oven malfunctions can manifest as abnormal heating, excessive noise during operation, and electrical leakage.
[0028] Fault diagnosis models can be trained diagnostic models.
[0029] Specifically, the training process for the fault diagnosis model can be performed as follows: First, collect and organize historical fault data of different household appliances. For each type of fault, organize corresponding fault voice description samples, fault appearance visual samples, fault diagnosis results, and corresponding repair solutions. Integrate all the organized data into a training dataset. Then, initialize the network parameters of the fault diagnosis model, input the training data into the initial model in batches, and perform forward propagation to obtain the predicted fault diagnosis results and repair solutions. Calculate the cross-entropy loss between the predicted results and the true labels, and use the gradient descent algorithm to update the model parameters in reverse. Repeat the above iterative process until the model loss drops to a preset threshold or reaches a preset number of iterations, and then terminate the training to obtain the final usable fault diagnosis model.
[0030] The fault diagnosis model can also access data from the diagnostic knowledge base and fault mode library.
[0031] Based on the experience of domain experts and historical fault data, a structured diagnostic knowledge base can be constructed, including fault types, fault causes, and repair solutions. Different types of household appliances correspond to different knowledge entries, facilitating rapid matching and retrieval by the model during the diagnostic process and improving the efficiency of generating diagnostic results. The fault pattern library summarizes the characteristic manifestations of various faults and standardizes the fault types corresponding to different users' colloquial descriptions, reducing diagnostic bias caused by non-standard user expressions and further ensuring the accuracy of fault diagnosis. After training, the fault diagnosis model can directly connect with the user-input text fault information and the collected visual information to quickly output the fault diagnosis result with the highest matching degree and the feasible repair solution. The knowledge base can be stored using a graph database, which facilitates rapid querying and reasoning; The knowledge base can include a sub-base of fault cause knowledge, a sub-base of fault solution knowledge, and a sub-base of fault performance standard terminology. Each sub-base is associated with a mapping between fault product type and fault code. When the model obtains the information provided by the user, it can quickly retrieve and match the corresponding knowledge content through the association relationship to support the generation of diagnostic results.
[0032] By collecting a large amount of fault data and extracting fault features, including vibration frequency, temperature changes, and current fluctuations, a fault mode library is constructed. The fault mode library can be classified and clustered using machine learning models such as support vector machines and random forests.
[0033] The fault mode library can include a set of standard fault modes under different product categories. Each fault mode corresponds to a unique mode code and is associated with a corresponding fault cause and solution index. It can quickly map user-speaking, non-standard fault descriptions into standard fault types, helping the fault diagnosis model to narrow the search scope and improve diagnostic efficiency and accuracy.
[0034] Step S2: Combine visual technology to determine the visual information of the product corresponding to the faulty product type.
[0035] Specifically, visual information can be obtained by combining visual technologies such as smart cameras to collect multi-angle images of the product corresponding to the type of faulty product.
[0036] Specifically, multiple appearance images of products corresponding to the faulty product type can be collected. After preprocessing each appearance image to remove background interference, the specific model of the faulty product can be identified through a pre-trained target detection algorithm. Visual information that can reflect fault characteristics in the product appearance can be extracted, such as shell damage, loose interfaces, smoke marks, water leakage marks, etc. This visual information is organized into structured data that the model can recognize for subsequent fault diagnosis.
[0037] Step S3: Convert the voice query signal into text to generate text fault information.
[0038] Specifically, the voice query signal can be converted into text by combining acoustic and language models to generate text fault information.
[0039] Specifically, automatic speech recognition technology can be used to segment the continuous speech signal input by the user into phoneme units, and then combine it with an acoustic model to calculate the corresponding text candidates. The language model is used to perform grammatical verification and semantic correction on the candidate results, and finally outputs fluent and accurate text fault information, retaining the core content such as the type of faulty product and the form of faulty manifestation described by the user, which is convenient for the fault diagnosis model to read and process.
[0040] The acoustic and language models, based on convolutional neural networks and recurrent neural networks, can adapt to user speech input with different accents and expression habits, effectively improving the accuracy of speech conversion and reducing the negative impact of speech recognition errors on subsequent diagnostic results. For user speech with regional accents, the model's generalization ability can accurately identify core fault information, ensuring stable speech conversion results for different user groups.
[0041] Text-based fault information contains text data corresponding to the type of faulty product and the form of fault manifestation. It can clearly carry the user's colloquial description of the fault, providing a standardized text input basis for subsequent fault diagnosis models and avoiding information noise interference caused by direct input of voice signals.
[0042] Step S4: Use the fault diagnosis model to perform fault diagnosis based on the text fault information and the visual information, and generate fault diagnosis results and repair solutions.
[0043] Specifically, the deep learning layers of the fault diagnosis model can be used to extract fault features from textual and visual fault information. The multi-task learning layers of the fault diagnosis model can be used to generate fault diagnosis results based on the fault features. The multi-objective optimization layers of the fault diagnosis model can be used to determine the repair scheme matched with the fault diagnosis results, ensuring the accuracy of the generated fault diagnosis results and repair schemes. It can integrate the dual features of textual and visual information, avoiding diagnostic bias caused by insufficient information from a single modality. Compared with diagnostic methods that rely solely on voice and text information, it can more accurately locate the fault location and cause, and output more targeted and executable repair schemes.
[0044] The reinforcement learning model simulates the effects of different repair solutions, selects the optimal solution, and further ensures the practicality and feasibility of the repair solution, avoiding the output of invalid solutions that do not match the current fault, thus saving users the time cost of screening and trying them themselves. The fault diagnosis model integrates user fault descriptions from textual fault information with appearance fault features from visual information, and can simultaneously combine the user's subjective description and the product's objective fault performance to complete fault location and cause analysis from two different dimensions, effectively reducing the probability of misjudgment.
[0045] By combining graph neural networks and knowledge bases, faults can be accurately located. By utilizing the relationship information between device components, the accuracy of the location can be improved. For example, when visual information identifies water leakage at the bottom of the rice cooker, combined with the text fault information described by the user's voice, such as "water is seeping out from the bottom when cooking rice and the heating is uneven," the model can quickly locate the fault as an aging seal ring fault by combining the connection relationship of the rice cooker components in the knowledge base, rather than generally judging it as a heating system fault. The output fault diagnosis results are more accurate.
[0046] Anomaly detection algorithms, such as Isolation Forest and Autoencoder, can be used to monitor equipment status in real time, identify abnormal behaviors, and combine the real-time monitored abnormal status data to further supplement the input information for fault diagnosis. This helps the fault diagnosis model to more comprehensively understand the current operating status of the equipment and improve the accuracy of fault location. For the monitored abnormal data, feature fusion can be performed with textual fault information and visual information, allowing the model to simultaneously combine user descriptions, appearance, and operating status information to complete the analysis, further narrowing down the possible scope of the fault and improving the reliability of the diagnostic results.
[0047] The system receives new fault data in real time and updates model parameters through online learning algorithms to adapt to changes in equipment status. During training, it uses labeled historical fault voice and text descriptions and equipment operation monitoring data as samples to continuously adjust model parameters, allowing the model to learn the data features corresponding to different fault types, thereby obtaining a stable and usable diagnostic model and ensuring the accuracy of the final output diagnostic results.
[0048] Transfer learning techniques can be used to transfer existing knowledge to new tasks, reducing data requirements.
[0049] Step S5: Optimize the fault diagnosis results and the repair plan using natural language processing technology, and generate text content.
[0050] Specifically, a Hidden Markov Model (HMM) or a neural network model incorporating natural language processing techniques can be used to remove non-linguistic components from the fault diagnosis results and repair solutions, segment lexical units and label them with parts of speech, perform case conversion and abbreviation expansion processing, and generate text content.
[0051] Step S6: Determine the voice features corresponding to the user.
[0052] Specifically, it can determine the user-defined voice features such as phonemes, pitch, speech rate, intonation, and volume. It can also extract user-specific voice features based on the user's historical voice input data. Furthermore, it can support users in independently adjusting and modifying voice feature parameters to meet the personalized voice output needs of different users.
[0053] This application can be integrated into other programs through API and SDK interfaces.
[0054] Among them, the API interface supports cross-platform calls and can be easily integrated into smart terminals and home management devices of different brands, meeting the secondary development needs of different hardware products.
[0055] The SDK interface provides a well-encapsulated development toolkit, eliminating the need for developers to build a voice feature processing module from scratch. Simple configuration is all that's required to integrate the functionality, reducing development costs and accelerating product launch. Through open interfaces, the fault diagnosis method presented in this application can flexibly adapt to various application scenarios. Whether it's a standalone intelligent voice fault diagnosis device or an additional functional module integrated into a smart home system, it can be seamlessly integrated and used, expanding the method's applicability. The acquired user voice features are encrypted and stored solely for generating personalized voice interaction signals for the user, ensuring the security of user information and complying with relevant personal information protection regulations.
[0056] Step S7: Based on the speech features, convert the text content into a speech interaction signal and play the speech interaction signal.
[0057] Specifically, text content can be converted into voice interaction signals by combining phoneme, pitch, and speech rate features, and then a vocoder can be used to play the voice interaction signals.
[0058] The generated voice interaction signals undergo post-processing, such as noise reduction and volume equalization, to improve clarity and audibility. The processed signals are then transmitted to an audio player or device for playback, or saved as audio files for later use. The system also comprehensively considers the emotional, prosodic, and personalized characteristics of the voice interaction signals. Specifically, during noise reduction, a deep learning model is used in conjunction with the emotional information of the voice interaction signals for adaptive noise reduction. For example, different noise reduction parameters and strategies are employed for voices expressing anger or anxiety to preserve subtle emotional nuances. Regarding volume equalization, the volume is dynamically adjusted based on the content, emotion, and user habits of the voice interaction signals; for instance, the volume is increased when emphasizing important information and decreased when expressing soothing emotions. Simultaneously, a prosodic feature adjustment module is introduced to automatically adjust the tone, speed, and pauses of the voice based on the semantics and emotion of the text, making the speech more natural and fluent. Furthermore, personalized processing of the voice signals can be performed based on user settings, such as accent and pronunciation habits, providing a more personalized voice interaction experience.
[0059] Simultaneously, this application also supports receiving secondary voice input from the user at any time during the playback of the voice interaction signal. Users can raise questions about the fault diagnosis results or repair solutions, or provide more details about the fault. This application can then re-analyze the fault according to the aforementioned process, quickly outputting updated diagnostic results and repair solutions. This meets the user's interactive needs for dynamically adjusting information and seeking in-depth consultation, avoiding deviations in diagnostic results due to incomplete initial information. It can gradually improve fault information through multiple rounds of interaction, continuously optimizing diagnostic results and providing users with more accurate fault diagnosis services. The entire interaction process is completed entirely by voice, eliminating the need for manual text input. This is particularly suitable for scenarios where both hands are busy handling faulty equipment and manual operation is inconvenient, providing users with a more convenient and seamless fault diagnosis experience.
[0060] As can be seen from the above technical solution, the fault diagnosis method based on intelligent voice provided in this application can acquire the user-inputted voice query signal and a trained fault diagnosis model. The voice query signal includes the type of faulty product and the manifestation of the fault. Therefore, in the process of using this application, only the user needs to input the voice query signal, eliminating the need for manual text input. This is especially convenient for people who are not proficient in text input or operating electronic devices, such as the elderly. For example, the elderly can initiate a query simply by stating the type of faulty household appliance and its manifestation without the effort of typing, saving time and simplifying the operation, thus greatly improving the efficiency of fault diagnosis. Simultaneously, this application can combine visual technology to determine the visual information of the product corresponding to the faulty product type; perform text conversion on the voice query signal to generate text fault information; and use the fault diagnosis model to perform fault diagnosis based on the text fault information and the visual information to generate fault diagnosis results and repair solutions. Based on this, this application can integrate multimodal information for fault diagnosis, improving the accuracy of fault diagnosis and repair solution determination. On the one hand, it utilizes the text fault information containing the faulty product type and fault manifestation after voice conversion; on the other hand, it combines the product visual information obtained through visual technology to provide data support for the fault diagnosis model from different dimensions, allowing the fault diagnosis model to analyze the fault more comprehensively and accurately, thereby matching the corresponding repair solutions. Subsequently, this application can use natural language processing technology to optimize the fault diagnosis results and the repair plan, and generate text content; determine the user's corresponding voice characteristics; convert the text content into a voice interaction signal based on the voice characteristics, and play the voice interaction signal; based on this, this application can use natural language processing technology to optimize the fault diagnosis results and repair plan, transforming professional and complex diagnosis and repair content into easy-to-understand text content, making it easier for users to understand; at the same time, the voice interaction signal played by this application is generated in combination with the user's corresponding voice characteristics, conforming to the user's voice habits, enhancing the interactive experience between the user and the system, especially for specific user groups such as the elderly, voice output is easier to accept and understand than text reading. Therefore, this application can integrate voice interaction, visual technology assistance, natural language optimization, and personalized voice output functions to provide fault diagnosis and repair plan determination services for the elderly who are not good at text search and operating electronic devices, as well as ordinary users who have difficulty understanding professional terminology, thus expanding the applicable population of fault diagnosis services. At the same time, it lowers the threshold for ordinary users to troubleshoot faults on their own. Common faults can be initially located and handled without the need for professional repair personnel to come to the site. It can help users quickly solve simple faults, saving time and economic costs for on-site repairs. It can also help professional repair personnel to make fault predictions in advance, prepare the corresponding tools and parts in advance, and improve the efficiency of on-site repairs.
[0061] In some embodiments of this application, the process of obtaining the user-input voice query signal in step S1 is described in detail, and the steps are as follows: S10. Respond to the user's voice input operation and perform intent recognition on the input voice.
[0062] Specifically, text analysis can be performed on the user's voice input to identify the user's intent.
[0063] Specifically, user voice can be captured via microphone. The captured voice is then preprocessed by a signal processing unit, including noise reduction, enhancement, and framing to remove background noise and signal interference, improving the quality of the voice signal and providing a clear and accurate data foundation for subsequent intent recognition. The preprocessed voice is then converted into text by a speech recognition engine. A pre-trained intent recognition classification model is used to match the user's intent. If the user's input is confirmed to be a fault diagnosis query, the voice query signal is extracted, containing the type of faulty product and the manifestation of the fault. If the user's intent is not a fault diagnosis query, the system redirects accordingly without executing the subsequent fault diagnosis process. This step filters user input through intent recognition, preventing irrelevant requests from consuming system resources and further improving the operational efficiency of the fault diagnosis service.
[0064] Among them, the speech recognition engine can use deep learning technology to convert speech signals into text information, thereby achieving efficient and accurate speech recognition. Speech recognition technology is one of the key technologies in this application. It uses acoustic models and language models to convert users' speech input into computer-readable text information. Currently, deep learning technologies, such as convolutional neural networks (CNN) and recurrent neural networks (RNN), have been widely used to improve the accuracy and robustness of speech recognition.
[0065] The noise reduction methods are as follows: Convolutional neural networks or recurrent neural networks are used to model the speech signal. By training on a large dataset of noisy and clean speech, the network can learn the characteristics of noise and effectively separate it. The filter parameters are dynamically adjusted by combining the least mean square error or recursive least squares algorithm to remove noise in real time. Microphone array technology is introduced to enhance the noise reduction effect through the spatial information of multi-channel signals. The enhancement method is as follows: a GAN model is used to generate high-quality speech signals; the generator network learns the mapping from noisy speech to clean speech, while the discriminator network distinguishes between generated speech and real clean speech; an attention mechanism is introduced into the speech enhancement model so that the model can focus on the speech harmonic structure of the speech signal; a multi-scale convolutional neural network is used to extract different frequency features of the speech signal and combine them with time-frequency domain information for enhancement.
[0066] S11. When it is determined that the input voice corresponds to a fault diagnosis, determine all data types contained in the input voice.
[0067] Specifically, it can be determined whether the user needs fault diagnosis based on the recognition intent; if so, it can determine all data types contained in the user's voice input.
[0068] For example, it can identify whether the user's voice input contains the form of the fault and the type of the faulty product.
[0069] If the user's voice input is missing some key data, such as only mentioning the type of faulty product without specifying the fault symptoms, or only describing the fault symptoms without specifying the type of faulty product, the system will automatically ask the user through voice interaction to fill in the missing information before proceeding to the next step, ensuring that subsequent fault diagnosis has complete and accurate data support.
[0070] S12. Match each data type with the preset fault diagnosis template to determine the supplementary data type. The fault diagnosis template is used to represent all type identifiers involved in the fault diagnosis.
[0071] Specifically, a fault diagnosis template can be preset, which contains all types of identifiers required for the voice query signal; Therefore, by matching the fault diagnosis template with various data types, the data types that need to be supplemented can be identified as supplementary data types.
[0072] S13. Generate interactive voice signals based on the supplementary data types.
[0073] Specifically, natural language processing technology can be used to convert supplementary data types into text supplementary suggestions; Text supplement suggestions can be converted into interactive voice signals based on speech characteristics.
[0074] Text supplement suggestions can be generated using natural and easy-to-understand conversational language based on the data type requirements, avoiding the use of technical and rigid expressions. This aligns with everyday communication habits, ensuring that the information needed is clearly communicated to the user without causing comprehension difficulties. For example, when the product type is missing, the suggestion could be set as "What device are you referring to that has a problem?", and when the symptoms are missing, it could be expressed as "Could you please describe the specific situation of the device now?", making the entire interaction process more natural and user-friendly.
[0075] Therefore, this application can proactively ask users questions to guide them to provide the missing supplementary information until all key data is supplemented, thus avoiding the inability to conduct diagnosis or the error of results due to missing information, and ensuring that the fault diagnosis process can proceed smoothly.
[0076] S14. Play the interactive voice signal to obtain the user's feedback voice signal.
[0077] Specifically, a vocoder can be used to play the interactive voice signal, guiding the user to provide additional feedback voice signals.
[0078] S15. Integrate the feedback voice signal and the input voice to obtain the voice query signal.
[0079] Specifically, the feedback voice signal and the input voice can be optimized and integrated to obtain a voice query signal. After the user provides supplementary feedback, the information of the corresponding supplementary data type in the feedback voice signal is extracted and integrated into the data of the original voice query signal. This ensures that the data input to the fault diagnosis model fully covers all necessary dimensions, laying the foundation for accurate diagnosis in the future.
[0080] As can be seen from the above technical solution, this embodiment provides an optional method for acquiring voice query signals. This method can confirm whether the user's input voice signal meets the fault diagnosis conditions when the user needs to perform fault diagnosis. If the fault diagnosis conditions are not met, the user is guided to supplement the information until the fault diagnosis conditions are met. This further improves the flexibility and applicability of this application.
[0081] In some embodiments of this application, the process of step S15, integrating the feedback voice signal and the input voice to obtain the voice query signal, is described in detail as follows: S150. Obtain a filter adjusted by the least mean square error or least squares algorithm, and use the filter to denoise the feedback speech signal and the input speech.
[0082] Specifically, convolutional neural networks or recurrent neural networks can be used to model speech signals. By training a large dataset of noisy and clean speech, the convolutional neural network or recurrent neural network can be trained so that it can learn the characteristics of noise and effectively separate noise. By combining the least mean square error or recursive least squares algorithm, the filter parameters are dynamically adjusted to remove noise in real time; microphone array technology is introduced to enhance the noise reduction effect through the spatial information of multi-channel signals. Filters, convolutional neural networks, or recurrent neural networks can be used to denoise the feedback speech signal and the input speech.
[0083] S151. By combining a speech enhancement model with an attention mechanism and a multi-scale convolutional speech network, speech enhancement is performed on the denoised feedback speech signal and the input speech.
[0084] Specifically, a GAN model can be used to generate high-quality speech signals; the generator network learns the mapping from noisy speech to clean speech, while the discriminator network distinguishes between generated speech and real clean speech; an attention mechanism is introduced into the speech enhancement model, enabling the model to focus on the speech harmonic structure of the speech signal; and a multi-scale convolutional neural network is used to extract different frequency features of the speech signal, which are then combined with time-frequency domain information for enhancement.
[0085] S152. The voice query signal is obtained by splicing the voice-enhanced feedback voice signal and the input voice.
[0086] Specifically, the feedback voice signal after noise reduction and voice enhancement can be integrated with the input voice to obtain a voice query signal. This facilitates subsequent unified text conversion and intent recognition, avoids data errors caused by processing in stages, ensures that the integrated signal is complete and clear, and can accurately reflect all the fault information provided by the user, providing a reliable data foundation for subsequent text conversion and fault diagnosis.
[0087] As can be seen from the above technical solution, this embodiment provides an optional method to integrate the feedback voice signal and the input voice to obtain the voice query signal. By combining the above method with denoising technology and voice enhancement technology, the reliability of the voice query signal of this application can be further improved, thereby ensuring the accuracy of data input in subsequent stages, reducing diagnostic errors caused by voice signal quality problems from the source, and further improving the accuracy of overall fault diagnosis. The specific process for training the fault diagnosis model is as follows: First, historical fault datasets of different product categories are collected. The datasets include labeled faulty product types, fault manifestations, root causes of faults, and corresponding repair solutions. The datasets are divided into training and testing sets according to a preset ratio. Then, the text fault information and corresponding visual information from the training set are input into the initial model. The error between the prediction result and the true label is calculated using the cross-entropy loss function. The backpropagation algorithm is used to continuously update the model's network parameters, gradually reducing the model's prediction error. During the training process, the model accuracy is periodically verified using the testing set, and the model's hyperparameters are adjusted to avoid overfitting until the model's diagnosis on the testing set is accurate. This ensures the accuracy of data input in subsequent stages, reduces diagnostic errors caused by speech signal quality issues at the source, and further improves the overall accuracy of fault diagnosis.
[0088] In some embodiments of this application, the process of obtaining the trained fault diagnosis model in step S1 is described in detail, and the steps are as follows: S10. Obtain training samples of different types of household appliances. Each training sample contains the corresponding household appliance's training diagnosis type, training fault manifestation form, training visual data, and training solution.
[0089] Specifically, it can generate training samples corresponding to different training and diagnostic types for different household appliances.
[0090] The same household appliance can correspond to different training fault manifestations, and the same training fault manifestation can correspond to different household appliances.
[0091] The various household appliances include, but are not limited to, refrigerators, washing machines, air conditioners, televisions, range hoods, water heaters, and other common household equipment, covering most common household appliance categories and meeting the daily fault diagnosis needs of ordinary users. The collected training samples are all from real fault cases of past official after-sales repairs, and the data is accurately and clearly labeled, which can provide real and reliable learning materials for model training, ensuring that the diagnostic logic learned by the model fits the actual repair scenario.
[0092] S11. Construct the initial diagnostic model.
[0093] Specifically, an initial diagnostic model can be constructed that can connect to a diagnostic knowledge base and a fault mode library.
[0094] The initial diagnostic model can include deep learning layers, multi-task learning layers, and multi-objective optimization layers.
[0095] S12. Train the initial diagnostic model using each training sample, and use the final initial diagnostic model as the trained fault diagnosis model.
[0096] Specifically, an uncertainty sampling mechanism can be introduced to train the initial diagnostic model using each training sample, and the final initial diagnostic model can be used as the trained fault diagnosis model.
[0097] During training, data samples with high model prediction uncertainty can be selected for manual annotation and feedback to update the model. Alternatively, a cluster-based sampling method can be used to identify different clusters in each training sample, and select representative training samples from each cluster to train the initial diagnostic model, ensuring that the model can learn diverse data features. It can also build meta-learning models, utilizing historical training data and optimization experience to quickly adapt to new scenarios and data distributions, thereby improving the model's generalization ability and learning efficiency.
[0098] As can be seen from the above technical solution, this embodiment provides an optional way to obtain a trained fault diagnosis model. The initial diagnosis model can be trained by combining training samples from multiple household appliances.
[0099] In some embodiments of this application, the process of step S10, obtaining training samples of different types of household appliances, wherein each training sample includes the training diagnosis type, training fault manifestation form, training visual data, and training solution for the corresponding household appliance, is described in detail as follows: The S100, combined with a smart camera, collects training visual data for different types of home appliances and different training and diagnostic types.
[0100] Specifically, images of different types of electrical appliances under different faults, captured by smart cameras, can be used as training visual data. Based on the fault and appliance type corresponding to each image, determine the training diagnosis type and household appliance corresponding to each training visual data.
[0101] S101. Using web crawler technology, obtain training fault manifestations and training solutions for different types of household appliances and different training diagnostic types.
[0102] Specifically, web crawling technology can be used to find the fault manifestations and solutions corresponding to different training and diagnostic types of different household appliances, which can be used as training fault manifestation forms and training solutions.
[0103] S102. For different household appliances, each training diagnostic type of the household appliance is integrated with the corresponding training visual data, training fault manifestations and training solutions to form a training sample.
[0104] Specifically, training visual data, training fault manifestations, training solutions, and the training diagnosis type corresponding to the same household appliance and the same training diagnosis type can be integrated to form training samples.
[0105] As can be seen from the above technical solution, this embodiment provides an optional method for obtaining training samples of different types of household appliances. The above method can further integrate smart camera and web crawler technology to improve data richness. At the same time, data is collected and integrated separately for different household appliances and different training and diagnostic types of the same appliance, making the training samples more targeted, thereby improving the reliability of this application.
[0106] In some embodiments of this application, the process of step S12, which involves training the initial diagnostic model using each training sample and using the resulting initial diagnostic model as the trained fault diagnosis model, is described in detail below: S120. Input each training sample into the initial diagnostic model in sequence to obtain the prediction scheme and prediction diagnosis result corresponding to each training sample.
[0107] Specifically, each training sample can be sequentially input into the initial diagnostic model to obtain the prediction scheme and prediction diagnosis result corresponding to each training sample output by the initial diagnostic model.
[0108] S121. Calculate the loss value of the initial diagnostic model based on each training sample, its corresponding prediction scheme, and the prediction diagnosis result.
[0109] Specifically, the training diagnosis type and training solution of each training sample can be compared with the corresponding prediction scheme and prediction diagnosis result by calculating the cosine similarity, and the cosine similarity can be used as the loss value.
[0110] S122. Using gradient descent, adjust the parameters of the initial diagnostic model based on the loss value.
[0111] Specifically, gradient descent can be used to adjust the parameters of the initial diagnostic model based on the loss value.
[0112] S123. Based on the loss value of each training sample, set the training frequency for each training sample.
[0113] Specifically, the training frequency of each training sample can be determined based on the magnitude of the loss value of each training sample.
[0114] The larger the loss value, the higher the corresponding training frequency can be.
[0115] S124. Train the initial diagnostic model with adjusted parameters based on each training sample and its corresponding training frequency.
[0116] Specifically, training samples can be selected and input into the initial diagnostic model by referring to the training frequency corresponding to each training sample, and the initial diagnostic model after parameter adjustment can be trained.
[0117] As can be seen from the above technical solution, this embodiment provides an optional method for training an initial diagnostic model. Through the above method, this application can further improve the training effect by combining gradient descent and uncertain training mechanisms.
[0118] In some embodiments of this application, the process of optimizing the fault diagnosis results and the repair plan using natural language processing technology and generating text content is described in detail as follows: S50. The fault diagnosis results and the repair plan are segmented into characters to obtain multiple fields.
[0119] Specifically, a conditional random field can be used to segment the fault diagnosis results and the repair scheme into multiple fields.
[0120] S51. Remove fields that are non-linguistic components and expand fields that are abbreviations to obtain multiple optimized fields.
[0121] Specifically, regular expressions can be used to match each field to remove non-linguistic elements from the text, such as punctuation marks, special characters, and numbers.
[0122] For certain complex non-linguistic components, such as abbreviations, slang, or special expressions, natural language understanding technology is needed for identification and processing, and additional contextual information or knowledge bases are required to assist in the judgment. By combining with the diagnostic pattern library, fields belonging to abbreviations can be expanded to obtain multiple optimized fields.
[0123] S52. Integrate the various optimized fields to generate text content.
[0124] Specifically, natural language processing technology can be used to integrate various optimized fields into text content.
[0125] As can be seen from the above technical solution, this embodiment provides an optional method for optimizing the fault diagnosis results and the repair plan using natural language technology and generating text content. This method can further improve the accuracy of the text content of this application and reduce the difficulty of user understanding.
[0126] Next, we will combine Figure 2 This application provides a detailed description of the intelligent voice-based fault diagnosis device. The intelligent voice-based fault diagnosis device described below can be compared with the intelligent voice-based fault diagnosis method described above.
[0127] See Figure 2 It can be observed that intelligent voice-based fault diagnosis devices may include: The acquisition module 10 is used to acquire the user-input voice query signal and the trained fault diagnosis model. The voice query signal includes the faulty product type and the fault manifestation. Combined with module 20, it is used to combine vision technology to determine the visual information of the product corresponding to the faulty product type; The conversion module 30 is used to convert the voice query signal into text and generate text fault information; The generation module 40 is used to perform fault diagnosis based on the text fault information and the visual information using the fault diagnosis model, and generate fault diagnosis results and repair solutions. The optimization module 50 is used to optimize the fault diagnosis results and the repair plan using natural language technology, and generate text content; Determining module 60 is used to determine the voice features corresponding to the user; The playback module 70 is used to convert the text content into a voice interaction signal based on the voice features, and to play the voice interaction signal.
[0128] Furthermore, the acquisition module 10 may include: The voice intent recognition unit is used to respond to the user's voice input operation and perform intent recognition on the input voice; A data type determination unit is used to determine all data types contained in the input voice when it is determined that the input voice corresponds to a fault diagnosis; The fault diagnosis template matching unit is used to match each data type with a preset fault diagnosis template to determine supplementary data types. The fault diagnosis template is used to characterize all types of identifiers involved in fault diagnosis. An interactive voice signal generation unit is used to generate interactive voice signals based on the supplementary data type. An interactive voice signal playback unit is used to play the interactive voice signal to obtain the user's feedback voice signal; A voice query signal generation unit is used to integrate the feedback voice signal and the input voice to obtain the voice query signal.
[0129] Furthermore, the voice query signal generation unit may include: The first voice query signal generation subunit is used to obtain a filter adjusted by the least mean square error or least squares algorithm, and to use the filter to denoise the feedback voice signal and the input voice. The second speech query signal generation subunit is used to combine a speech enhancement model with an attention mechanism and a multi-scale convolutional speech network to enhance the denoised feedback speech signal and the input speech. The third voice query signal generation subunit is used to splice the voice-enhanced feedback voice signal and the input voice to obtain the voice query signal.
[0130] Furthermore, the acquisition module 10 may also include: The training sample acquisition unit is used to acquire training samples of different types of household appliances. Each training sample contains the corresponding household appliance's training diagnosis type, training fault manifestation form, training visual data, and training solution. The initial diagnostic model building unit is used to build the initial diagnostic model; An initial diagnostic model training unit is used to train the initial diagnostic model using each training sample, and the final initial diagnostic model is used as the trained fault diagnosis model.
[0131] Furthermore, the training sample acquisition unit may include: The first training sample acquisition subunit is used to combine with smart cameras to collect training visual data of different types of home appliances and different training and diagnostic types. The second training sample acquisition subunit is used to acquire training fault manifestations and training solutions for different types of household appliances and different training diagnostic types through web crawling technology. The third training sample acquisition subunit is used to integrate each training diagnostic type of the household appliance with the corresponding training visual data, training fault manifestations, and training solutions to form training samples for different household appliances.
[0132] Furthermore, the initial diagnostic model training unit may include: The first initial diagnostic model training subunit is used to sequentially input each training sample into the initial diagnostic model to obtain the prediction scheme and prediction diagnosis result corresponding to each training sample. The second initial diagnostic model training subunit is used to calculate the loss value of the initial diagnostic model based on each training sample and its corresponding prediction scheme and prediction diagnosis result. The third initial diagnostic model training subunit is used to adjust the parameters of the initial diagnostic model based on the loss value using the gradient descent method. The fourth initial diagnostic model training subunit is used to set the training frequency of each training sample based on the loss value of each training sample. The fifth initial diagnostic model training subunit is used to train the parameter-adjusted initial diagnostic model based on each training sample and its corresponding training frequency.
[0133] Furthermore, the optimization module 50 may include: A character segmentation unit is used to segment the fault diagnosis results and the repair plan into multiple fields. The optimized field generation unit is used to remove fields that are non-linguistic components and expand fields that are abbreviations to obtain multiple optimized fields. The optimized field integration unit is used to integrate various optimized fields to generate text content.
[0134] The intelligent voice-based fault diagnosis device provided in this application embodiment can be applied to intelligent voice-based fault diagnosis equipment, such as PC terminals, cloud platforms, servers, and server clusters. Optionally, Figure 3 The hardware structure block diagram of the fault diagnosis device based on intelligent voice is shown. (Refer to...) Figure 3 The hardware structure of a fault diagnosis device based on intelligent voice may include: at least one processor 1, at least one communication interface 2, at least one memory 3 and at least one communication bus 4. In this embodiment of the application, the number of processor 1, communication interface 2, memory 3, and communication bus 4 is at least one, and processor 1, communication interface 2, and memory 3 communicate with each other through communication bus 4; Processor 1 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. Memory 3 may include high-speed RAM, and may also include non-volatile memory, such as at least one disk storage device; The memory stores a program, which the processor can call. The program is used for: The system acquires user-input voice query signals and a trained fault diagnosis model, wherein the voice query signals include the type of faulty product and the form of fault manifestation. By combining visual technology, the visual information of the product corresponding to the type of faulty product is determined; The voice query signal is converted into text to generate text fault information; The fault diagnosis model is used to perform fault diagnosis based on the text fault information and the visual information, and to generate fault diagnosis results and repair solutions. The fault diagnosis results and the repair plan are optimized using natural language processing technology, and text content is generated. Determine the voice features corresponding to the user; Based on the speech features, the text content is converted into a speech interaction signal, and the speech interaction signal is played.
[0135] Optionally, the refined and extended functions of the program can be referred to the above description.
[0136] This application embodiment also provides a readable storage medium that can store a program suitable for execution by a processor, the program being used for: The system acquires user-input voice query signals and a trained fault diagnosis model, wherein the voice query signals include the type of faulty product and the form of fault manifestation. By combining visual technology, the visual information of the product corresponding to the type of faulty product is determined; The voice query signal is converted into text to generate text fault information; The fault diagnosis model is used to perform fault diagnosis based on the text fault information and the visual information, and to generate fault diagnosis results and repair solutions. The fault diagnosis results and the repair plan are optimized using natural language processing technology, and text content is generated. Determine the voice features corresponding to the user; Based on the speech features, the text content is converted into a speech interaction signal, and the speech interaction signal is played.
[0137] Optionally, the refined and extended functions of the program can be referred to the above description.
[0138] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0139] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0140] In the several embodiments provided in this application, it should be understood that the disclosed devices, electronic devices, computer storage media, computer program products, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units and modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0141] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0142] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0143] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0144] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. The various embodiments of this application can be combined with each other. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A fault diagnosis method based on intelligent voice, characterized in that, include: The system acquires user-input voice query signals and a trained fault diagnosis model, wherein the voice query signals include the type of faulty product and the form of fault manifestation. By combining visual technology, the visual information of the product corresponding to the type of faulty product is determined; The voice query signal is converted into text to generate text fault information; The fault diagnosis model is used to perform fault diagnosis based on the text fault information and the visual information, and to generate fault diagnosis results and repair solutions. The fault diagnosis results and the repair plan are optimized using natural language processing technology, and text content is generated. Determine the voice features corresponding to the user; Based on the speech features, the text content is converted into a speech interaction signal, and the speech interaction signal is played.
2. The fault diagnosis method based on intelligent voice according to claim 1, characterized in that, Acquire the user's voice query signal, including: Respond to the user's voice input and perform intent recognition on the input voice; When it is determined that the input voice corresponds to a fault diagnosis, all data types contained in the input voice are determined; Each data type is matched with a preset fault diagnosis template to determine supplementary data types. The fault diagnosis template is used to characterize all types of identifiers involved in fault diagnosis. Based on the aforementioned supplementary data types, an interactive voice signal is generated; Play the interactive voice signal to obtain the user's feedback voice signal; The feedback voice signal and the input voice signal are integrated to obtain the voice query signal.
3. The fault diagnosis method based on intelligent voice according to claim 2, characterized in that, The process of integrating the feedback voice signal and the input voice to obtain the voice query signal includes: Obtain a filter adjusted by the least mean square error or least squares algorithm, and use the filter to denoise the feedback speech signal and the input speech; By combining a speech enhancement model with an attention mechanism and a multi-scale convolutional speech network, speech enhancement is performed on the denoised feedback speech signal and the input speech. The voice query signal is obtained by splicing the enhanced feedback voice signal and the input voice.
4. The fault diagnosis method based on intelligent voice according to claim 1, characterized in that, Obtain the trained fault diagnosis model, including: Acquire training samples of different types of household appliances. Each training sample contains the corresponding household appliance's training diagnosis type, training fault manifestation form, training visual data, and training solution. Construct an initial diagnostic model; The initial diagnostic model is trained using each training sample, and the final initial diagnostic model is used as the trained fault diagnosis model.
5. The fault diagnosis method based on intelligent voice according to claim 4, characterized in that, The acquisition of training samples for different types of household appliances includes: Combine smart cameras to collect training visual data for different types of home appliances and different training and diagnostic types; By using web crawling technology, we can obtain the training fault manifestations and training solutions for different types of household appliances and different training diagnostic types. For each type of household appliance, each training diagnostic type of the household appliance is integrated with the corresponding training visual data, training fault manifestations, and training solutions to form a training sample.
6. The fault diagnosis method based on intelligent voice according to claim 4, characterized in that, The step of training the initial diagnostic model using each training sample includes: Each training sample is sequentially input into the initial diagnostic model to obtain the prediction scheme and prediction diagnosis result corresponding to each training sample; Based on each training sample and its corresponding prediction scheme and prediction diagnosis result, the loss value of the initial diagnosis model is calculated; Based on the loss value, the parameters of the initial diagnostic model are adjusted using the gradient descent method. Based on the loss value of each training sample, set the training frequency for each training sample; The initial diagnostic model, after parameter adjustment, is trained based on each training sample and its corresponding training frequency.
7. The fault diagnosis method based on intelligent voice according to claim 1, characterized in that, The process of optimizing the fault diagnosis results and the repair plan using natural language processing technology and generating text content includes: The fault diagnosis results and the repair plan are segmented into characters to obtain multiple fields; Fields that are non-linguistic components are removed, and fields that are abbreviations are expanded to obtain multiple optimized fields; The various optimized fields are integrated to generate text content.
8. A fault diagnosis device based on intelligent voice, characterized in that, include: The acquisition module is used to acquire the user-input voice query signal and the trained fault diagnosis model. The voice query signal includes the faulty product type and the fault manifestation. The module is used to combine visual technology to determine the visual information of the product corresponding to the faulty product type. The conversion module is used to convert the voice query signal into text and generate text fault information; The generation module is used to perform fault diagnosis based on the text fault information and the visual information using the fault diagnosis model, and generate fault diagnosis results and repair solutions. An optimization module is used to optimize the fault diagnosis results and the repair plan using natural language processing technology, and to generate text content. A determination module is used to determine the voice features corresponding to the user; The playback module is used to convert the text content into a voice interaction signal based on the voice features, and to play the voice interaction signal.
9. A fault diagnosis device based on intelligent voice, characterized in that, Including memory and processor; The memory is used to store programs; The processor is used to execute the program to implement each step of the fault diagnosis method based on intelligent voice as described in any one of claims 1-7.
10. A readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements each step of the fault diagnosis method based on intelligent voice as described in any one of claims 1-7.