AI Server Multi-Modal Home Appliance Failure Diagnosis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for determining the failure type of home appliances are inaccurate and require numerous questions and answers, as they rely solely on language or text, which struggles to convey the appliance's state effectively.
Innovation Solution
An artificial intelligence server that uses a multi-modal learning method to classify failure types by processing both image and sound data from a video of the home appliance, employing a classification model that combines feature vectors from image and voice classification models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If language or text is used to identify failure type, then the service can be provided remotely, but the accuracy of conveying the appliance state is insufficient
Solution Approach 1:
The patent combines multiple data modalities (image data from video, sound data from video, and text data from user input) into a unified analysis framework. The multi-modal learning model integrates these different types of data to comprehensively determine failure types, thereby improving identification accuracy while maintaining remote service capability.
2Device complexity
If only image data is used for classification, then the processing is simple, but the failure type determination is inaccurate
Solution Approach 1:
The patent employs a composite multi-modal learning model that integrates multiple types of data (image, sound, and text) rather than relying on a single data type. This composite approach allows the system to leverage the strengths of each modality, improving failure type determination accuracy while keeping the overall system architecture manageable through modular design.
3Loss of information
If multiple questions and answers are exchanged to identify failure, then more information can be gathered, but the time and operational steps increase
Solution Approach 1:
The system performs preliminary analysis by automatically extracting and processing image and sound data from the uploaded video before user interaction. This preliminary action allows the system to pre-process and understand the appliance state, reducing the need for extensive back-and-forth questioning and thereby shortening the overall diagnosis time while maintaining information completeness.
Data Source
AI summary
An artificial intelligence server includes: a communication interface configured to communicate with a terminal, and a processor. The processor is configured to receive, from the terminal through the communication interface, a video of a home appliance, acquire a first feature vector by inputting image data extracted from the video to an image classification model, acquire a second feature vector by inputting sound data extracted from the video to a voice classification model, acquire a result value by inputting a data set obtained by combining the first feature vector and the second feature vector to an abnormality classification model, and transmit, to the terminal through the communication interface, a failure type acquired based on the result value.


