Maintenance guidance method and device for household appliances
By constructing a retrieval and reasoning framework for multimodal feature fusion and a closed-loop self-optimization mechanism, the problems of strong skill dependence and information loss in household appliance fault diagnosis are solved, and high-accuracy fault diagnosis and dynamic knowledge updating are achieved.
Patent Information
- Application Number
- CN202511172341.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-08-21
AI Technical Summary
Existing household appliance fault diagnosis and repair rely on the personal experience of maintenance technicians, and have problems such as strong skill dependence, inefficient information retrieval, loss of information dimensions and rigid knowledge system, resulting in low fault identification accuracy and repair efficiency.
Build a retrieval and reasoning framework based on multimodal feature fusion, directly analyze the image, audio and text features of the fault through a unified vector database, generate maintenance guidance plans based on a large language model, and dynamically update the knowledge base through a closed-loop self-optimization mechanism.
The overall accuracy of fault diagnosis has been improved to over 95%, avoiding information dimensionality reduction loss, reducing the risk of knowledge hallucination in large language models in professional fields, and achieving dynamic updating and real-time performance of the knowledge base.
Smart Images

Figure CN120670561A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of household appliances, and in particular to a maintenance guidance method and device for household appliances. Background Art
[0002] At present, fault diagnosis and repair in the kitchen appliance industry mainly rely on the personal experience of repair technicians, and the following defects are common: (1) Strong skill dependence: The skill levels of repair personnel vary, especially for new technicians, resulting in low fault identification accuracy and repair efficiency. (2) Inefficient information retrieval: Traditional knowledge bases (such as manuals and documents) are mainly text-based, and it is difficult to quickly and accurately locate the complex audio and video fault information on site. (3) Loss of information dimension: When processing fault pictures or videos, existing technical solutions often convert them into text descriptions first. This process will lose a lot of key visual (such as burn marks on components, tiny cracks) and auditory (such as the frequency and rhythm of specific abnormal sounds) details, which are the core basis for accurate diagnosis. (4) Solidified knowledge system: Traditional knowledge bases are static and cannot absorb new cases and new techniques generated during the front-line repair process. The knowledge system will gradually become outdated and difficult to cope with the endless new problems. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a maintenance guidance method and device for household appliances, so as to construct a retrieval and reasoning framework based on the direct fusion of multimodal features. By directly analyzing the image, audio and text features of the fault, the information dimensionality loss is avoided, and the comprehensive accuracy of fault diagnosis is increased to more than 95%; through dynamic knowledge retrieval and enhanced generation technology, the risk of "knowledge hallucination" that may occur in general large language models in professional fields is effectively reduced.
[0004] In the first aspect, an embodiment of the present invention provides a maintenance guidance method for household appliances, which pre-constructs a unified vector database based on text data, image data and audio data related to household appliance failures. The method includes: obtaining on-site situation information of the household appliance, the on-site situation information including: video information, audio information and / or text information; performing image matching, audio matching and / or text matching on the on-site situation information based on the unified vector database, and outputting knowledge fragments as retrieval results; inputting the retrieval results into a large language model, and outputting a maintenance guidance plan for the household appliance.
[0005] In an optional embodiment of the present application, the above method also includes: extracting text vectors corresponding to text materials through a large text model, extracting image vectors corresponding to image materials through a visual feature extraction model, and extracting audio vectors corresponding to audio materials through an audio recognition model; and constructing a unified vector database based on text vectors, image vectors and audio vectors.
[0006] In an optional embodiment of the present application, the step of extracting text vectors corresponding to text data through a large text model includes: semantically slicing the text data through the large text model to obtain text vectors; the step of extracting image vectors corresponding to image data through a visual feature extraction model includes: encoding the visual information of the image data as an image vector through a visual feature determination model; the step of extracting audio vectors corresponding to audio data through an audio recognition model includes: extracting spectral features of the audio data through the audio recognition model to obtain audio vectors corresponding to the audio data.
[0007] In an optional embodiment of the present application, the above-mentioned unified vector database includes: text vectors, image vectors and audio vectors, text data, image data and audio data, and metadata; metadata includes: device model, component name, data source link, and associated fault label.
[0008] In an optional embodiment of the present application, the above-mentioned steps of performing image matching, audio matching and text matching on the scene situation information based on the unified vector database, and outputting knowledge fragments as retrieval results, include: determining the query image vector, query audio vector and query text vector based on the scene situation information; performing image matching on the query image vector based on the unified vector database, and determining the most similar visual scene as the knowledge fragment; performing audio matching on the query audio vector based on the unified vector database, and determining the most similar sound as the knowledge fragment; performing text matching on the query text vector based on the unified vector database, and determining the most similar problem description as the knowledge fragment; determining the association weights of multiple knowledge fragments, and sorting the multiple knowledge fragments based on the association weights to obtain retrieval results.
[0009] In an optional embodiment of the present application, the above-mentioned step of sorting multiple knowledge fragments based on association weights includes: determining the signal quality of multiple knowledge fragments; adjusting the association weights of multiple knowledge fragments based on the signal quality; and sorting the multiple knowledge fragments based on the adjusted association weights.
[0010] In an optional embodiment of the present application, the above-mentioned step of inputting the retrieval results into the large language model and outputting the maintenance guidance plan for the household appliance includes: inputting the retrieval results into the large language model and outputting the maintenance text guidance for the household appliance; injecting the knowledge fragment into the Prompt template; inputting the Prompt template into the large language model and outputting the structured maintenance guidance; wherein the information elements of the Prompt template include: text knowledge information, key visual evidence information, key auditory evidence information and information for answering user questions; and using the maintenance text guidance and the structured maintenance guidance as the maintenance guidance plan for the household appliance.
[0011] In an optional embodiment of the present application, the above method also includes: obtaining a maintenance report on the completion of maintenance of the household appliance; the maintenance report includes structured feedback and unstructured feedback, the structured feedback is used for the maintenance technician to confirm maintenance-related issues in the form of multiple-choice questions, and the unstructured feedback is used for the maintenance technician to enter text through a text box; determining the target case marked by the maintenance technician based on the maintenance report; wherein the target case represents a case that has been successfully solved and has a large difference from the existing unified vector database; adding the on-site situation information of the target case for feature vectorization and the maintenance guidance plan of the target case to the unified vector database; adjusting the association weight in the unified vector database based on the maintenance report; if the maintenance report represents that the target maintenance guidance plan is an incorrect maintenance guidance plan, the knowledge fragment associated with the target maintenance guidance plan in the unified vector database is marked with a label representing low credibility.
[0012] In an optional embodiment of the present application, the above-mentioned text large model generates structured feedback based on the video information, audio information and / or text information involved in this maintenance; the text large model evaluates the results of this maintenance based on the evaluation of the maintenance technician, and performs model learning based on the evaluation results.
[0013] In a second aspect, an embodiment of the present invention further provides a maintenance guidance device for household appliances, which pre-constructs a unified vector database based on text data, image data and audio data related to household appliance failures. The device includes: a retrieval result output module, which is used to obtain on-site situation information of household appliances, and the on-site situation information includes: video information, audio information and / or text information; based on the unified vector database, image matching, audio matching and / or text matching are performed on the on-site situation information respectively, and knowledge fragments are output as retrieval results; a maintenance guidance plan generation module, which is used to input the retrieval results into a large language model and output a maintenance guidance plan for household appliances.
[0014] The embodiments of the present invention bring the following beneficial effects: An embodiment of the present invention provides a method and device for providing maintenance guidance for household appliances. The method pre-builds a unified vector database based on text, image, and audio data related to household appliance faults; obtains on-site information about the household appliance, including video, audio, and / or text information; performs image matching, audio matching, and / or text matching on the on-site information based on the unified vector database, and outputs knowledge fragments as retrieval results; inputs the retrieval results into a large language model, and outputs a maintenance guidance plan for the household appliance. In this method, a retrieval and reasoning framework based on direct fusion of multimodal features can be constructed. By directly analyzing the image, audio, and text features of the fault, information dimensionality reduction losses are avoided, and the overall accuracy of fault diagnosis is increased to over 95%. Dynamic knowledge retrieval and enhanced generation techniques are used to effectively reduce the risk of "knowledge hallucination" that may occur in general large language models in professional fields.
[0015] Other features and advantages of the present disclosure will be set forth in the following description, or some features and advantages may be inferred or unambiguously determined from the description, or may be learned by practicing the above-mentioned technology of the present disclosure.
[0016] In order to make the above-mentioned objectives, features and advantages of the present disclosure more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0018] Figure 1 A schematic diagram of an offline end constructing a unified vector database according to an embodiment of the present invention; Figure 2 A flowchart of a household appliance maintenance guidance method provided by an embodiment of the present invention; Figure 3 A schematic diagram of real-time online diagnosis and interaction provided by an embodiment of the present invention; Figure 4 A flowchart of another household appliance maintenance guidance method provided by an embodiment of the present invention; Figure 5 A schematic diagram of a closed-loop self-optimization mechanism provided by an embodiment of the present invention; Figure 6A schematic diagram of updating association weights in a unified vector database and marking labels representing low credibility provided by an embodiment of the present invention; Figure 7 A schematic structural diagram of a maintenance guidance device for household appliances provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] At present, fault diagnosis and repair in the kitchen appliance industry mainly rely on the personal experience of repair technicians, and the following defects are common: (1) Strong skill dependence: The skill levels of repair personnel vary, especially for new technicians, resulting in low fault identification accuracy and repair efficiency. (2) Inefficient information retrieval: Traditional knowledge bases (such as manuals and documents) are mainly text-based, and it is difficult to quickly and accurately locate the complex audio and video fault information on site. (3) Loss of information dimension: When processing fault pictures or videos, existing technical solutions often convert them into text descriptions first. This process will lose a lot of key visual (such as burn marks on components, tiny cracks) and auditory (such as the frequency and rhythm of specific abnormal sounds) details, which are the core basis for accurate diagnosis. (4) Solidified knowledge system: Traditional knowledge bases are static and cannot absorb new cases and new techniques generated during the front-line repair process. The knowledge system will gradually become outdated and difficult to cope with the endless new problems.
[0021] Based on this, the embodiments of the present invention provide a household appliance maintenance guidance method and device, which specifically provides a kitchen appliance fault diagnosis and maintenance guidance system based on multimodal feature fusion and closed-loop self-optimization, which can avoid information dimensionality reduction loss and improve the overall accuracy of fault diagnosis.
[0022] To facilitate understanding of this embodiment, a maintenance guidance method for a household appliance disclosed in an embodiment of the present invention is first introduced in detail.
[0023] Example 1: The embodiment of the present invention provides a maintenance guidance method for household appliances, which constructs a unified vector database based on text data, image data and audio data related to household appliance failures. Figure 1 The diagram shown is a schematic diagram of an offline end constructing a unified vector database. In this embodiment, a multimodal knowledge base, namely a unified vector database, can be constructed based on original data such as text data, image data and audio data.
[0024] In some embodiments, text vectors corresponding to text materials can be extracted through a large text model, image vectors corresponding to image materials can be extracted through a visual feature extraction model, and audio vectors corresponding to audio materials can be extracted through an audio recognition model; a unified vector database can be constructed based on text vectors, image vectors, and audio vectors.
[0025] In this embodiment, multimodal feature vectorization can be performed to convert unstructured raw data into machine-understandable, high-fidelity feature vectors.
[0026] In some embodiments, the text data can be semantically sliced through a large text model to obtain a text vector; the visual information of the image data can be encoded as an image vector through a visual feature determination model; and the spectral features of the audio data can be extracted through an audio recognition model to obtain an audio vector corresponding to the audio data.
[0027] 1. Text data processing: For PDF (Portable Document Format) manuals, SOP (Standard Operating Procedure) documents, etc., use large text models such as DeepSeek-V3 to perform semantic slicing and generate text embedding vectors.
[0028] 2. Image data processing: For fault images and video key frames, use visual feature extraction models such as CLIP (Contrastive Language–Image Pre-training) and ViT (Vision Transformer) to generate image embedding vectors. This vector can encode visual information such as image color, texture, shape, and component status.
[0029] 3. Video data processing: For the audio tracks in the maintenance videos or individual fault recordings, use audio recognition models such as Audio Spectrogram Transformer to extract their spectral features and generate audio embedding vectors. These vectors can distinguish different types of mechanical friction sounds, electric current sounds, or gas leak sounds.
[0030] In some embodiments, the unified vector database includes: text vectors, image vectors and audio vectors, text data, image data and audio data, and metadata; metadata includes: device model, component name, data source link, and associated fault label.
[0031] like Figure 1As shown, this embodiment can construct a unified vector database based on text vectors, image vectors, and audio vectors. This unified vector database can be built using professional vector databases that support multi-vector indexing, such as Milvus and Weaviate. This database can simultaneously store: three types of embedding vectors (text, image, and audio); raw data (text snippets, image files, and audio / video clips); and metadata (such as device models, component names, data source links, and associated fault tags).
[0032] Based on the above description, see Figure 2 The flowchart of a maintenance guidance method for a household appliance is shown, and the maintenance guidance method for a household appliance includes the following steps: Step S202: Acquire on-site situation information of household appliances, which includes video information, audio information and / or text information; perform image matching, audio matching and / or text matching on the on-site situation information based on a unified vector database, and output knowledge fragments as retrieval results.
[0033] See also Figure 3 The diagram shows a real-time online diagnosis and interaction system. A technician can use a mobile app (application) to collect on-site information about household appliances. This information can include short videos, images, and audio / text descriptions. This embodiment uses a unified vector database to perform image matching, audio matching, and / or text matching on this information, outputting the corresponding knowledge fragments as search results.
[0034] In some embodiments, the query image vector, query audio vector and query text vector can be determined based on the on-site situation information; image matching is performed on the query image vector based on the unified vector database to determine the most similar visual scene as the knowledge fragment; audio matching is performed on the query audio vector based on the unified vector database to determine the most similar sound as the knowledge fragment; text matching is performed on the query text vector based on the unified vector database to determine the most similar problem description as the knowledge fragment; the association weights of multiple knowledge fragments are determined, and the multiple knowledge fragments are sorted based on the association weights to obtain retrieval results.
[0035] This embodiment can determine the query image vector, query audio vector and query text vector based on the on-site situation information. For example, key frame images are extracted from the video to generate a query image vector; audio is extracted from the video to generate a query audio vector; and the technician's voice or text description is used to generate a query text vector.
[0036] like Figure 3As shown, this embodiment can also perform cross-modal fusion retrieval, that is, perform parallel, cross-modal similarity searches. For example: audio matching: using the query audio vector to search for the most similar sound in the audio vectors of the knowledge base; image matching: using the query image vector to search for the most similar visual scene in the image vectors of the knowledge base; text matching: using the query text vector to search for the most similar question description in the text vectors of the knowledge base.
[0037] In this embodiment, a dynamic weighted fusion algorithm can be used to sort the three-way search results. For example, if the energy and characteristics of the input audio signal are significant (such as a strong abnormal sound), the weight of the audio matching result will be dynamically increased.
[0038] In some embodiments, signal qualities of the plurality of knowledge fragments may be determined; association weights of the plurality of knowledge fragments may be adjusted based on the signal qualities; and the plurality of knowledge fragments may be ranked based on the adjusted association weights.
[0039] The dynamic weighted fusion algorithm can be: ; The dynamic weight adjustment rules can be:
[0040] Aq, Adb: Aq represents the query audio vector, which is the vector converted from the audio signal input by the user for retrieval; Adb is the audio vector in the knowledge base, which is used for similarity matching with the query audio vector.
[0041] Iq, Idb: Iq is the query image vector, that is, the vector converted from the image input by the user for retrieval; Idb is the image vector in the knowledge base, which is used to calculate the similarity with the query image vector.
[0042] Tq, Tdb: Tq refers to the query text vector, which is the vector converted from the text input by the user for retrieval; Tdb is the text vector in the knowledge base, which is used for similarity matching with the query text vector.
[0043] Sim(Aq, Adb), Sim(Iq, Idb), Sim(Tq, Tdb): are the similarities between the query audio and the knowledge base audio, the query image and the knowledge base image, and the query text and the knowledge base text, respectively, used to measure the degree of matching between the two.
[0044] α, β, and γ are dynamically adjusted weight coefficients, corresponding to the weights of audio, image, and text matching results in the fusion score, respectively. They change dynamically based on signal quality (such as audio energy, image clarity, and text confidence).
[0045] Audio_Energy: The peak value of the input audio spectrum energy (dB), reflecting the audio energy strength. When the energy is high (such as strong abnormal sound), the audio matching weight will be dynamically increased.
[0046] Image_Clarity_Score: A score based on image blur / contrast (0-1). The higher the score, the clearer the image is. This affects the weight of the image matching result during fusion.
[0047] Text_Confidence: The confidence level of keyword matching in text descriptions. A high confidence level indicates a good match between the text and the knowledge base text, which will affect the weight of the text matching results.
[0048] Fused_Score: fusion score, calculated by combining the similarity results of audio, image, and text matching with their respective weights α, β, and γ through a dynamic weighted fusion algorithm, and used to sort search results.
[0049] Max_Energy: This is the maximum reference value of the audio energy, which is used to compare with the actual Audio_Energy to assist in determining the dynamic adjustment of the audio weight α.
[0050] Threshold: Threshold. When Audio_Energy exceeds this threshold, the weight α of the audio matching result will be dynamically adjusted according to the Audio_Energy / Max_Energy method.
[0051] In this embodiment, the weights of audio, image, and text matching results in the fusion score can be dynamically changed based on signal quality (such as audio energy, image clarity, text confidence, etc.), and the associated weights can be flexibly adjusted to accurately sort multiple knowledge fragments.
[0052] Step S204: input the search results into the large language model and output a maintenance guidance plan for the household appliance.
[0053] In some embodiments, the retrieval results can be input into a large language model to output text instructions for repairing household appliances; knowledge fragments can be injected into a Prompt template; the Prompt template can be input into the large language model to output structured repair instructions; wherein, the information elements of the Prompt template include: text knowledge information, key visual evidence information, key auditory evidence information and information for answering user questions; the text instructions for repairing household appliances and the structured repair instructions are used as repair guidance plans for household appliances.
[0054] like Figure 3As shown, this embodiment can also perform enhanced reasoning and generation based on retrieval results, inputting the retrieval results into a large language model to output textual repair instructions for household appliances. Using the RAG (Retrieval-Augmented Generation) architecture, multiple retrieved, top-scoring multimodal knowledge fragments (perhaps a text description, a best-matching fault image, or a most similar abnormal noise audio clip) are injected into a domain-optimized Prompt template (a structured framework for specific output types).
[0055] For example, a prompt could read: "You are an experienced kitchen appliance repair expert. Please combine the following high-fidelity knowledge: [text knowledge: {{retrieved_text}}], [key visual evidence: {{retrieved_image_link}}], and [key auditory evidence: {{retrieved_audio_link}}] to answer the user's question: {{query}}."
[0056] like Figure 3 As shown, this embodiment also allows for system question-and-answering. Prompts are fed into a large language model, such as Qwen2.5-32b, to generate structured repair steps. The answers not only include text instructions but also embed links to the most relevant images and video clips retrieved from the knowledge base, allowing technicians to view them at any time.
[0057] An embodiment of the present invention provides a method for providing maintenance guidance for household appliances. The method pre-builds a unified vector database based on text, image, and audio data related to household appliance faults; obtains on-site information about the household appliance, including video, audio, and / or text; performs image matching, audio matching, and / or text matching on the on-site information based on the unified vector database, and outputs knowledge fragments as retrieval results; inputs the retrieval results into a large language model, and outputs a maintenance guidance plan for the household appliance. In this method, a retrieval and reasoning framework based on direct fusion of multimodal features can be constructed. By directly analyzing the image, audio, and text features of the fault, information dimensionality reduction losses are avoided, and the overall accuracy of fault diagnosis is increased to over 95%. Dynamic knowledge retrieval and enhanced generation techniques are used to effectively reduce the risk of "knowledge hallucination" that may occur in general large language models in professional fields.
[0058] Example 2: This embodiment provides another maintenance guidance method for household appliances. This method is implemented on the basis of the above embodiment, focusing on the specific method of the closed-loop self-optimization mechanism. Figure 4 A flowchart of another method for providing maintenance guidance for a household appliance is shown, and the method for providing maintenance guidance for a household appliance comprises the following steps: Step S402: Acquire on-site situation information of household appliances, which includes video information, audio information and / or text information; perform image matching, audio matching and / or text matching on the on-site situation information based on a unified vector database, and output knowledge fragments as retrieval results.
[0059] Step S404: input the search results into the large language model and output a maintenance guidance plan for the household appliance.
[0060] Step S406: Obtain a maintenance report of the household appliance; perform feature vectorization on the maintenance report and add it to the unified vector database.
[0061] See also Figure 5 The diagram shows a closed-loop self-optimization mechanism. In this embodiment, feature vectorization can be performed on the maintenance report of the household appliance and the report can be added to the unified vector database.
[0062] The above maintenance report includes structured feedback and unstructured feedback. The structured feedback is used for maintenance technicians to confirm maintenance-related issues in the form of multiple-choice questions, and the unstructured feedback is used for maintenance technicians to enter text in a text box.
[0063] The structured feedback can also be generated by the large text model based on the video, audio, and / or text information involved in the repair. The large text model can evaluate the repair results based on the technician's evaluation and conduct a new round of model learning based on the evaluation results.
[0064] For example, structured feedback involves multiple-choice questions to confirm the root cause of the fault, the replaced parts, and whether the repair was successful. Unstructured feedback involves a text box where the technician can evaluate the effectiveness of the system's guidance or provide additional information on new techniques or issues discovered during the repair.
[0065] In some embodiments, the target case marked by the maintenance technician can be determined based on the maintenance report; wherein the target case represents a case that has been successfully resolved and has a large difference from the existing unified vector database; the on-site situation information of the target case to be feature vectorized and the maintenance guidance plan of the target case are added to the unified vector database.
[0066] like Figure 5 As shown, this embodiment can automatically store knowledge: for target cases marked as "successfully solved" by maintenance technicians and with significant differences from the existing knowledge base, the system will treat them (including the initial multimodal input and the solution confirmed by the technician) as a high-quality Q&A pair, automatically perform feature vectorization, and add them to the unified knowledge base after a small amount of manual review.
[0067] In some embodiments, the association weights in the unified vector database may also be adjusted based on the maintenance report.
[0068] See also Figure 6 The diagram below shows how to update association weights and label low-confidence tags in a unified vector database. This embodiment also allows for dynamic adjustment of knowledge weights: the system statistically analyzes feedback data. If a solution is frequently adopted and receives positive reviews, its association weight and search ranking in the knowledge base will automatically increase. Conversely, if a solution is frequently marked as "invalid," its weight will be reduced, achieving the survival of the fittest in knowledge.
[0069] In some embodiments, if the maintenance report indicates that the target maintenance guidance solution is an erroneous maintenance guidance solution, the knowledge segment associated with the target maintenance guidance solution in the unified vector database is labeled with a label indicating low credibility.
[0070] like Figure 6 As shown, this embodiment can also perform self-correction of "knowledge illusion": for guidance plans generated by large models but proven to be wrong in front-line practice, the system can automatically label its associated knowledge fragments as "low credibility" based on negative feedback, thereby suppressing its appearance in future searches.
[0071] The above method provided by the embodiment of the present invention can establish a closed-loop self-optimizing knowledge base iteration mechanism, and use the real-time feedback data of front-line maintenance personnel to enable the system knowledge base to self-learn, self-correct and continuously grow, forming a dynamically evolving "living" knowledge system.
[0072] This embodiment also provides a specific example, which is explained by taking "range hood does not smoke" as an example.
[0073] 1. Technician input: Take a video of the motor running, which contains a clear metal friction sound, and verbally describe "the motor is running, but there feels no suction."
[0074] 2. System Multimodal Analysis: Generate an audio vector (encoding the characteristics of metal friction sound), generate an image vector (encoding the appearance of the motor and the rotation state of the turbine), and generate a text vector (from the speech-to-text conversion of "The motor is running, but there is no suction").
[0075] 3. Cross-modal fusion search results: Audio match: Highly similar to the audio sample "turbine bearing failure or foreign object" in the knowledge base (similarity 0.91). Text match: Correlated with the text description "motor idling, insufficient suction" (similarity 0.85). Image match: Matches the image of "turbine rotating normally," ruling out a stuck turbine. Fusion ranking: Due to the highly significant audio features, the audio match result receives the highest weight, and the system highly suspects a bearing or foreign object issue.
[0076] 4. Generate Guidance Plan: (1) Preliminary diagnosis: Based on the sound and images you provided, it is highly suspected that the internal bearing of the turbine is damaged or foreign matter has been sucked in, causing friction.
[0077] (2) Repair steps: Safe operation: Always disconnect the power supply to the device. Disassembly inspection: Remove the turbine air box (reference video: V-2035, click to view disassembly details). Problem location: Check the turbine bearing for looseness or metal debris (compare to image: I-5012, bearing damage characteristics), and check for foreign objects in the air duct. Solution: If bearing damage is found, the entire turbine module needs to be replaced (Note: Please check the device model when replacing...).
[0078] 5. Closed-Loop Feedback: After the technician completes the repair, he selects "Fault Cause: Bearing Damage" in the app and clicks "Repair Successful." This case is then learned by the system, increasing the correlation between "metal friction sound" and "bearing damage."
[0079] In summary, the core of the above-mentioned household appliance maintenance guidance method provided by the embodiment of the present invention is that, offline, the system no longer forcibly converts all maintenance data into text. Instead, it uses multiple AI models to extract multimodal feature vectors from the equipment manual (text), fault diagram (image), and maintenance video (video keyframe + audio), respectively, to build a unified, multi-dimensional vector knowledge base. During online diagnosis, the system directly analyzes the fault video captured on-site, extracts the feature vectors of its image, audio, and text descriptions, performs cross-modal fusion retrieval, and accurately matches the most similar fault scenarios (including similar images, similar sounds, and similar descriptions) in the knowledge base. Finally, using the RAG architecture, the retrieved high-fidelity multimodal knowledge is injected into the large language model to generate an accurate maintenance guidance plan and attach a link to the most relevant original data (pictures, video clips). At the same time, the present invention designs a maintenance feedback loop to feed back successful front-line cases into the knowledge base, realizing system self-learning and iteration.
[0080] The above method provided by the embodiment of the present invention can construct a retrieval and reasoning framework in the field of kitchen appliances based on the direct fusion of multimodal features. By directly analyzing the image, audio and text features of the fault, it can avoid information dimensionality reduction loss and improve the overall accuracy of fault diagnosis. It can also establish a closed-loop self-optimization knowledge base iteration mechanism in the field of kitchen appliances, and use the real-time feedback data of front-line maintenance personnel to enable the system knowledge base to self-learn and self-correct, dynamically "evolve" the knowledge system, and enhance the real-time and accuracy of knowledge.
[0081] Example 3: Corresponding to the above method embodiment, the embodiment of the present invention provides a maintenance guidance device for household appliances, which pre-builds a unified vector database based on text data, image data and audio data related to household appliance failures. Figure 7 The structure diagram of a maintenance guidance device for a household appliance shown in FIG. 1 is a schematic diagram of a maintenance guidance device for a household appliance, wherein the maintenance guidance device for a household appliance comprises: A search result output module 71 is configured to obtain on-site information about household appliances, including video information, audio information, and / or text information; perform image matching, audio matching, and / or text matching on the on-site information based on a unified vector database, and output a knowledge fragment as a search result; The maintenance guidance solution generating module 72 is used to input the search results into the large language model and output the maintenance guidance solution for the household appliance.
[0082] An embodiment of the present invention provides a maintenance guidance device for household appliances. The device pre-builds a unified vector database based on text, image, and audio data related to household appliance faults; obtains on-site information about the household appliance, including video, audio, and / or text information; performs image matching, audio matching, and / or text matching on the on-site information based on the unified vector database, and outputs knowledge fragments as retrieval results; inputs the retrieval results into a large language model, and outputs a maintenance guidance plan for the household appliance. In this method, a retrieval and reasoning framework based on direct fusion of multimodal features can be constructed. By directly analyzing the image, audio, and text features of the fault, information dimensionality reduction losses are avoided, and the overall accuracy of fault diagnosis is increased to over 95%. Dynamic knowledge retrieval and enhanced generation technology are used to effectively reduce the risk of "knowledge illusion" that may occur in general large language models in professional fields.
[0083] The above-mentioned unified vector database construction module is used to extract text vectors corresponding to text materials through a large text model, extract image vectors corresponding to image materials through a visual feature extraction model, and extract audio vectors corresponding to audio materials through an audio recognition model; and construct a unified vector database based on text vectors, image vectors and audio vectors.
[0084] The unified vector database construction module is used to perform semantic slicing on text materials through a large text model to obtain text vectors; to determine the visual information of the image materials encoded by the visual feature model as image vectors; and to extract the spectral features of the audio materials through the audio recognition model to obtain the audio vectors corresponding to the audio materials.
[0085] The unified vector database includes: text vectors, image vectors and audio vectors, text data, image data and audio data, and metadata; metadata includes: device model, component name, data source link, and associated fault label.
[0086] The above-mentioned retrieval result output module is used to determine the query image vector, query audio vector and query text vector based on the on-site situation information; perform image matching on the query image vector based on the unified vector database to determine the most similar visual scene as the knowledge fragment; perform audio matching on the query audio vector based on the unified vector database to determine the most similar sound as the knowledge fragment; perform text matching on the query text vector based on the unified vector database to determine the most similar problem description as the knowledge fragment; determine the association weights of multiple knowledge fragments, sort the multiple knowledge fragments based on the association weights, and obtain the retrieval results.
[0087] The retrieval result output module is used to determine the signal quality of multiple knowledge fragments; adjust the association weights of the multiple knowledge fragments based on the signal quality; and sort the multiple knowledge fragments based on the adjusted association weights.
[0088] The above-mentioned maintenance guidance solution generation module is used to input the search results into the large language model and output text maintenance guidance for household appliances; inject knowledge fragments into the Prompt template; input the Prompt template into the large language model and output structured maintenance guidance; wherein the information elements of the Prompt template include: text knowledge information, key visual evidence information, key auditory evidence information and information for answering user questions; and use the maintenance text guidance and structured maintenance guidance as the maintenance guidance solution for household appliances.
[0089] The above-mentioned device also includes: a unified vector database update module, which is used to obtain a maintenance report on the completion of maintenance of household appliances; the maintenance report includes structured feedback and unstructured feedback, the structured feedback is used for the maintenance technician to confirm maintenance-related issues in the form of multiple-choice questions, and the unstructured feedback is used for the maintenance technician to enter text through a text box; based on the maintenance report, the target case marked by the maintenance technician is determined; wherein, the target case represents a case that is successfully solved and has a large difference from the existing unified vector database; the on-site situation information of the target case for feature vectorization and the maintenance guidance plan of the target case are added to the unified vector database; the association weight in the unified vector database is adjusted based on the maintenance report; if the maintenance report represents that the target maintenance guidance plan is an incorrect maintenance guidance plan, the knowledge fragment associated with the target maintenance guidance plan in the unified vector database is marked with a label representing low credibility.
[0090] The unified vector database update module is also used by the text model to generate structured feedback based on the video information, audio information and / or text information involved in this maintenance; the text model evaluates the results of this maintenance based on the maintenance technician's evaluation and performs model learning based on the evaluation results.
[0091] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the household appliance maintenance guidance device described above can refer to the corresponding process in the aforementioned embodiment of the household appliance maintenance guidance method, and will not be repeated here.
[0092] In addition, in the description of the embodiments of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances.
[0093] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0094] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and therefore should not be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0095] Finally, it should be noted that the above embodiments are only specific implementation methods of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that any person skilled in the art can modify or easily conceive of changes to the technical solutions described in the above embodiments within the technical scope disclosed by the present invention, or replace some of the technical features therein with equivalents. Such modifications, changes or replacements do not deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.
Claims
1. A maintenance guidance method for household appliances, characterized in that: A unified vector database is constructed in advance based on text data, image data, and audio data related to household appliance failures. The method includes: Acquiring on-site situation information of the household appliance, the on-site situation information including: video information, audio information and / or text information; performing image matching, audio matching and / or text matching on the on-site situation information based on the unified vector database, and outputting a knowledge fragment as a retrieval result; The search results are input into a large language model, and a maintenance guidance plan for the household appliance is output.
2. The method according to claim 1, characterized in that The method further comprises: The text vector corresponding to the text data is extracted through the text large model, the image vector corresponding to the image data is extracted through the visual feature extraction model, and the audio vector corresponding to the audio data is extracted through the audio recognition model; A unified vector database is constructed based on the text vector, the image vector, and the audio vector.
3. The method according to claim 2, characterized in that The step of extracting text vectors corresponding to text data using the large text model includes: semantically slicing the text data using the large text model to obtain text vectors; The step of extracting an image vector corresponding to the image data by using a visual feature extraction model includes: encoding visual information of the image data as an image vector by using a visual feature determination model; The step of extracting an audio vector corresponding to the audio data through the audio recognition model includes: extracting the frequency spectrum features of the audio data through the audio recognition model to obtain the audio vector corresponding to the audio data.
4. The method according to claim 2, characterized in that The unified vector database includes: the text vector, the image vector and the audio vector, the text material, the image material and the audio material, and metadata; The metadata includes: device model, component name, data source link, and associated fault label.
5. The method according to claim 1, wherein The steps of performing image matching, audio matching, and text matching on the on-site situation information based on the unified vector database and outputting knowledge fragments as retrieval results include: Determine a query image vector, a query audio vector, and a query text vector based on the scene situation information; performing image matching on the query image vector based on the unified vector database to determine the most similar visual scene as the knowledge fragment; Performing audio matching on the query audio vector based on the unified vector database to determine the most similar sound as the knowledge fragment; Performing text matching on the query text vector based on the unified vector database to determine the most similar question description as the knowledge fragment; Determine the association weights of the plurality of knowledge fragments, sort the plurality of knowledge fragments based on the association weights, and obtain a retrieval result.
6. The method according to claim 5, characterized in that The step of sorting the plurality of knowledge fragments based on the association weights comprises: determining signal qualities of a plurality of said knowledge fragments; adjusting the association weights of the plurality of the knowledge fragments based on the signal quality; The plurality of knowledge fragments are sorted based on the adjusted association weights.
7. The method according to claim 5, characterized in that The step of inputting the search results into a large language model and outputting a maintenance guidance solution for the household appliance comprises: Inputting the search results into a large language model and outputting text instructions for repairing the household appliance; Injecting the knowledge fragment into a prompt template; inputting the prompt template into a large language model to output structured maintenance instructions; wherein the information elements of the prompt template include: textual knowledge information, key visual evidence information, key auditory evidence information, and information answering user questions; The maintenance text guidance and the structured maintenance guidance are used as a maintenance guidance solution for the household appliance.
8. The method according to any one of claims 1 to 7, characterized in that The method further comprises: Obtaining a maintenance report for the household appliance; the maintenance report includes structured feedback and unstructured feedback, wherein the structured feedback is used to allow the maintenance technician to confirm maintenance-related issues in the form of multiple-choice questions, and the unstructured feedback is used to allow the maintenance technician to enter text in a text box; Determining a target case marked by a maintenance technician based on the maintenance report; wherein the target case represents a case that has been successfully resolved and has a significant difference from an existing unified vector database; Adding the on-site situation information of the target case subjected to feature vectorization and the maintenance guidance plan of the target case to the unified vector database; adjusting the association weights in the unified vector database based on the maintenance report; If the maintenance report indicates that the target maintenance guidance solution is an erroneous maintenance guidance solution, a label indicating low credibility is marked on the knowledge segment associated with the target maintenance guidance solution in the unified vector database.
9. The method according to claim 8, characterized in that The method further comprises: The text model generates the structured feedback based on the video information, audio information and / or text information involved in the maintenance; The large text model evaluates the results of this maintenance based on the maintenance technician's evaluation and conducts model learning based on the evaluation results.
10. A maintenance guidance device for household appliances, characterized in that: A unified vector database is constructed in advance based on text data, image data, and audio data related to household appliance failures. The device includes: a retrieval result output module, configured to obtain on-site information of household appliances, the on-site information including video information, audio information, and / or text information; perform image matching, audio matching, and / or text matching on the on-site information based on the unified vector database, and output a knowledge fragment as a retrieval result; The maintenance guidance solution generating module is used to input the search results into a large language model and output a maintenance guidance solution for the household appliance.
Citation Information
Patent Citations
Multimedia file retrieval method
CN109271533A
Video fuzzy search method and system based on big data
CN114943009A
Household appliance fault diagnosis method, terminal device and server
CN119312207A
Personalized information search system for different types of barrier groups
CN119884409A
Zero-training vehicle re-identification method based on visual language large model
CN120220085A
Cited By
Equipment abnormity visual diagnosis method, electronic equipment and storage medium
CN122090358A
Device abnormal visual diagnosis method, electronic device and storage medium
CN122090358B