Train fault detection method and device, electronic equipment and storage medium
By adopting a combination of preset rules and multimodal large models in train fault detection, the problems of low accuracy and efficiency in complex fault identification in existing technologies are solved, and efficient identification and safety assurance of complex faults are achieved.
Patent Information
- Application Number
- CN202411487747.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-23
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2044-10-23
AI Technical Summary
Existing train fault detection methods are inaccurate and inefficient when dealing with complex faults, and lack multi-model integrated decision-making capabilities, making it difficult to effectively identify complex faults.
Preset rules are used to process image inference results with the first attribute type, and the multimodal large model is used to process image inference results with the second attribute type. The multimodal large model is used to assist in in-depth analysis to ensure accurate identification of complex faults.
It improves the accuracy and efficiency of fault detection, enhances the adaptability and processing depth of complex faults, and provides more comprehensive protection for the safe operation of trains.
Smart Images

Figure CN119559121B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of industrial quality inspection, in particular to the technical field of vehicles, and especially to a train fault detection method and device, an electronic device, and a storage medium. BACKGROUND
[0002] In the industrial field, defect detection is a crucial link in the entire production process. Improving product quality to produce high-value-added and high-profit products can achieve a leap in product competitiveness. The ideas for improving product quality include strengthening quality inspection, improving process technology standards, and standardizing production operations, among which strengthening quality inspection is the most commonly used way in manufacturing production.
[0003] Taking the railway train fault detection scene as an example, there are many types of faults to be detected in the railway train fault detection scene, including a large number of complex faults that require common sense understanding to distinguish. However, such fault data is small and usually cannot be distinguished by relying on local features. Therefore, how to successfully distinguish various faults is an urgent problem to be solved. SUMMARY
[0004] The present disclosure provides a train fault detection method, device, electronic device, and storage medium.
[0005] According to an aspect of the present disclosure, a train fault detection method is provided, comprising:
[0006] obtaining images taken by each detection station in the train, and obtaining corresponding image inference results according to the taken images, wherein the image inference results contain prediction results corresponding to the taken images;
[0007] if the attribute type of the target image to be detected is the first type, processing the image inference results according to a preset rule to obtain a detection result corresponding to the target image to be detected;
[0008] if the attribute type of the target image to be detected is the second type, simultaneously processing the image inference results by using a multi-modal large model to obtain a fault result corresponding to the target image to be detected.
[0009] According to another aspect of the present disclosure, a train fault detection device is provided, comprising:
[0010] an obtaining module configured to obtain images taken by each detection station in the train, and obtain corresponding image inference results according to the taken images, wherein the image inference results contain prediction results corresponding to the taken images;
[0011] The first processing module is configured to, if the attribute type of the to-be-detected target image is a first type, process the image inference result according to a preset rule to obtain a detection result corresponding to the to-be-detected target image.
[0012] The second processing module is configured to, if the attribute type of the to-be-detected target image is a second type, simultaneously process the image inference result by using a multi-modal large model to obtain a fault result corresponding to the to-be-detected target image.
[0013] According to a third aspect of the present disclosure, an electronic device is provided, comprising:
[0014] at least one processor; and
[0015] a memory in communication with the at least one processor; wherein
[0016] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of the above technical solutions.
[0017] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to perform the method of any one of the above technical solutions.
[0018] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising a computer program which, when executed by a processor, implements the method of any one of the above technical solutions.
[0019] The present disclosure provides a train fault detection method, device, equipment and storage medium. When detecting a to-be-detected target image with a first type of attribute, the present disclosure processes an image inference result according to a preset rule. When detecting a to-be-detected target image with a second type of attribute, the present disclosure simultaneously analyzes the to-be-detected target image by using a multi-modal large model after using the same processing method as the first type, and obtains a detection result corresponding to the to-be-detected target image by using the multi-modal large model. The to-be-detected target fault with the second type of attribute is analyzed in depth by using the multi-modal large model, so that complex faults can be accurately identified. Therefore, this method can improve the accuracy and efficiency of fault detection, and also enhances the adaptability and processing depth for complex faults, thereby providing more comprehensive protection for the safe operation of trains.
[0020] It should be appreciated that the content described in this section is not intended to identify key or critical features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0021] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them:
[0022] Figure 1 is a schematic diagram of the steps of the train fault detection method in an embodiment of the present disclosure;
[0023] Figure 2 is a schematic diagram of the fault detection process for the attribute type of the target image to be detected being the first type in an embodiment of the present disclosure;
[0024] Figure 3 is a schematic diagram of the fault detection process for the attribute type of the target image to be detected being the second type in an embodiment of the present disclosure;
[0025] Figure 4 is a schematic diagram of the fault detection process for the target image to be detected in an embodiment of the present disclosure;
[0026] Figure 5 is a schematic diagram of the principle block diagram of the train fault detection device in an embodiment of the present disclosure;
[0027] Figure 6 is a block diagram of an electronic device for implementing the train fault detection method in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0028] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to help understanding, and should be considered as merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, in order to be clear and concise, descriptions of well-known functions and structures are omitted in the following description.
[0029] The existing train fault detection method can be mainly divided into the following several ways: the first way is to detect faults by artificial means, that is, to detect each part that needs to be detected by artificial means; the second way is to detect faults based on traditional image processing, and the specific way is: to process the pictures collected by the camera through traditional image processing algorithm, to quantify the defects, and to detect the defects meeting the conditions by setting threshold; the third way is to detect faults based on general deep learning technology.
[0030] However, the above methods have the following disadvantages: for the first method, artificial defect detection requires workers to be at fixed workstations for a long time, and to observe the product by naked eye to determine whether there is a defect. In this way, not only is the labor intensity great, but the detection stability and consistency are also poor. In addition, the artificial quality inspection method generally uses paper and pen to record the quality inspection results, which leads to incomplete and scattered data, and cannot form valuable feedback information to guide lean production. For the second method, the traditional image processing-based defect detection method generally needs to be designed for specific products, and a series of rules are defined according to shape, color, width, height, etc. to perform defect detection. This method has poor robustness, and all rules and algorithms need to be redesigned and developed as the detected object changes, and even a set of rules cannot be reused for different batches of the same product. For the third method, the existing defect detection method based on deep learning technology generally directly reuses the general image object detection and segmentation algorithm, and does not use existing prior knowledge to design specific algorithms for the specific scene of defect detection, so there is still a lot of room for improvement in effect. In addition, the existing deep learning scheme basically has a single model design post-processing process and does not have the ability to make comprehensive decisions with multiple models.
[0031] To solve the above technical problems, the present disclosure provides a new train fault detection method, as shown in Figure 1 The method comprises the following steps:
[0032] In step S101, images taken by each detection station in the train are obtained, and corresponding image inference results are obtained according to the taken images. The image inference results contain prediction results corresponding to the taken images.
[0033] Specifically, when detecting the fault of the train, multiple detection stations are generally used to take pictures of the car bottom and the car body side, so that different station images can be obtained. After obtaining the images taken by different stations, the images taken by different stations are generally input into different target detection models for inference and prediction to obtain image inference results corresponding to the images taken by different stations. The image inference results contain prediction results corresponding to the taken images. The prediction results can be the category of the object contained in the taken images, and can also be other key information contained in the taken images. In this way, by obtaining the image inference results corresponding to the images taken by each detection station in advance, it is beneficial to subsequent fault judgment.
[0034] In step S102, if the attribute type of the target image to be detected is the first type, the image inference results are processed according to the preset rules to obtain the detection results corresponding to the target image to be detected.
[0035] Specifically, the to-be-detected target image is an image that needs to be judged, and whether a fault that needs to be judged exists in the image, where the number of to-be-detected target images is multiple. After obtaining the image inference result, different fault detection modes can be determined according to the type of the to-be-detected target image, so the attribute type of the to-be-detected target image can be judged before different fault detection modes are determined. If the attribute type of the to-be-detected target image is a first type, a preset rule is used to perform simple and controllable explicit logical discrimination on the image inference result, so that the detection result corresponding to the to-be-detected target image can be obtained. The detection result corresponding to the to-be-detected target image may be, for example, that a fault exists or that no fault exists.
[0036] The preset rule can be a post-processing operation that is pre-set, and can include specific rules such as name matching and threshold filtering. Referring to Table 1 below, the specific preset rule can include the content in Table 1.
[0037]
[0038] Table 1
[0039] In this way, the image inference result is processed by the preset rule, so that the detection result corresponding to the to-be-detected target image can be obtained. In addition, the correspondence between the preset rule and the image inference result, that is, which image inference result is processed by which preset rule, is pre-set, and a person skilled in the art can set it according to actual needs, which is not limited herein.
[0040] In step S103, if the attribute type of the to-be-detected target image is a second type, the image inference result is processed by a multi-modal large model at the same time, and the fault result corresponding to the to-be-detected target image is obtained by the multi-modal large model.
[0041] Specifically, it needs to be explained that the first type and the second type have different attributes. Specifically, the first type can be considered as a simple type, and the second type is a complex type. The division of the simple type and the complex type can be specifically divided according to actual application. For example, the first type and the second type are pre-divided, as long as a specified component structure object exists in the train, it is considered as a complex type, that is, the second type, and other component structures are considered as a simple type, that is, the first type. Therefore, the attribute type of the to-be-detected target image can be obtained by judging whether a specified component structure object exists in the to-be-detected image.
[0042] In the embodiments of the present disclosure, if the attribute type of the to-be-detected target image is the second type, the image inference result needs to be processed by means of the multi-modal large model, that is, the same processing mode as the first type is first adopted, and then the multi-modal large model is used to obtain the fault result corresponding to the to-be-detected target image. The detection result corresponding to the to-be-detected target image may be, for example, that there is a fault, or that there is no fault.
[0043] The present disclosure provides a train fault detection method, device, equipment and storage medium. When detecting a to-be-detected target image with an attribute type of a first type, the present disclosure processes an image inference result according to a preset rule. When detecting a to-be-detected target image with an attribute type of a second type, the present disclosure analyzes the to-be-detected target image by means of a multi-modal large model after adopting the same processing mode as the first type, and obtains a detection result corresponding to the to-be-detected target image by means of the multi-modal large model. The to-be-detected target fault with the attribute type of the second type is analyzed in depth by means of the multi-modal large model, so as to ensure that complex faults can be accurately identified. As can be seen, this method not only improves the accuracy and efficiency of fault detection, but also enhances the adaptability and processing depth for complex faults, thereby providing more comprehensive protection for the safe operation of trains.
[0044] In some optional embodiments, images photographed by each detection station in the train are obtained, and corresponding image inference results are obtained according to the photographed images, including:
[0045] Obtaining images photographed by each detection station in the train;
[0046] Inputting the images photographed by each detection station into different target detection models respectively, and obtaining image inference results corresponding to the images photographed by each detection station by means of the target detection models.
[0047] Specifically, when detecting the fault of the train, a plurality of detection stations are generally used to photograph the underframe and the side of the vehicle body to obtain photographed images of different stations. After obtaining the photographed images of different stations, the photographed images of different stations are generally input into different models for inference prediction to obtain image inference results corresponding to the images of different stations. The image inference result contains a prediction result corresponding to the photographed image. The prediction result may be the category of the object contained in the photographed image, and may also be other key information contained in the photographed image.
[0048] In this way, the image inference results corresponding to the photographed images of each detection station are obtained in advance, which is beneficial to subsequent fault judgment.
[0049] In some optional embodiments, if the attribute type of the to-be-detected target image is the first type, the image inference result is processed according to a preset rule to obtain a detection result corresponding to the to-be-detected target image, including:
[0050] If the attribute type of the to-be-detected target image is the first type, the business logic contained in the to-be-detected target image is analyzed to obtain a business analysis result.
[0051] The image inference result to be used in the image inference result is determined according to the business analysis result to obtain a first target image inference result.
[0052] The first target image inference result is processed according to a preset rule to obtain a detection result corresponding to the to-be-detected target image.
[0053] Specifically, if the attribute type of the to-be-detected target image is the first type, the business logic contained in the to-be-detected target image is analyzed to obtain a business analysis result, and then the first target image inference result is determined according to the business analysis result. For the business logic contained in the to-be-detected target image, it is used to indicate which station image inference result needs to be used together to judge whether the to-be-detected target image exists. For example, taking detecting whether a burglar device in the to-be-detected target image is damaged as an example, by analyzing the to-be-detected target image, it can be obtained that the business logic contained in the to-be-detected target image is that the burglar device is tied with a wire, and then the burglar device is damaged. At this time, it can be known that the image inference result obtained by inputting the to-be-detected target image into the wire detection model and the part positioning detection model is needed, and then the two image inference results are used as the first target image inference result. After obtaining the first target image inference result, the first target image inference result is processed according to a preset rule, and finally a detection result corresponding to the to-be-detected target image can be obtained.
[0054] In this way, by analyzing the business logic contained in the to-be-detected target image, it is beneficial to determine which image inference result needs to be processed, that is, it is beneficial to determine the first target image inference result, and then it is beneficial to improve the fault detection efficiency.
[0055] In some optional embodiments, the first target image inference result is processed according to a preset rule to obtain a detection result corresponding to the to-be-detected target image, including:
[0056] If the first target image inference result includes at least two image inference results, each image inference result in the first target image inference result is processed according to a preset rule to obtain a processed image inference result.
[0057] The processed image inference results are logically operated to obtain a detection result corresponding to the target image to be detected.
[0058] Specifically, after obtaining the target image inference result, it is necessary to determine which image inference results are included in the first target image inference result. If at least two image inference results are included in the first target image inference result, the preset rule is used to process each image inference result in the target image inference result respectively to obtain a processed image inference result, and then the processed image inference results are logically operated. Referring to Table 2, Table 2 shows that the logical operation includes logical AND, logical OR and logical NOT.
[0059] Table 2
[0060] After obtaining the processed image processing result, the processed image processing results are logically operated, thereby facilitating obtaining a detection result corresponding to the target image to be detected. If the target image inference result only contains one image inference result, the detection result corresponding to the target image to be detected can be directly obtained according to the image inference result.
[0061] In this way, different logical operation processes are performed on the image inference results according to the number of image inference results included in the target image inference result, thereby facilitating accurate determination of whether a fault exists and improving the accuracy of fault judgment.
[0062] In order to facilitate the scheme of the embodiments of the present application, the following examples are given, referring to Figure 2 , Figure 2 is a fault detection flowchart of the target image to be detected in the embodiment of the present disclosure. Taking whether a security device damage fault exists in the target image to be detected as an example, by analyzing the target image to be detected, it can be obtained that the business logic contained in the target image to be detected is that the security device is bound with a wire, and then the security device is damaged. At this time, it can be known that the image inference results obtained by using the wire detection model and the part positioning detection model are needed, and the wire model inference result and the part positioning inference result are obtained first. After obtaining the wire model inference result and the part positioning inference result, the two image inference results are taken as target image inference results. Since the target image inference result includes two image inference results, the two image inference results are respectively subjected to name matching, threshold filtering, NMS (non-maximum suppression), and IOU (intersection over union) calculation operations to obtain respective processed image inference results, which are wire and security device respectively. Next, the two processed image inference results are logically AND processed. Since the wire and the security device exist simultaneously, it can be obtained that the target image to be detected exists, that is, the security device is damaged.
[0063] In some optional embodiments, if the attribute type of the target image to be detected is the second type, the image inference result is processed by the multi-modal large model at the same time, and the fault result corresponding to the target image to be detected is obtained through the multi-modal large model, including:
[0064] If the attribute type of the target image to be detected is the second type, after the same processing mode as the first type is adopted, the second target image inference result is determined according to the business logic contained in the target image to be detected, and the prompt word is generated;
[0065] The prompt word and the second target image inference result are input into the multi-modal large model, and the fault result corresponding to the target image to be detected is obtained through the multi-modal large model.
[0066] Specifically, if the attribute type of the target image to be detected is the second type, after the same processing mode as the first type is adopted, the image inference result needs to be processed by the multi-modal large model at the same time, specifically, the second target image inference result is determined according to the business logic contained in the target image to be detected, and the prompt word is generated, and then the prompt word and the second target image inference result are input into the multi-modal large model, and the fault result corresponding to the target image to be detected is obtained through the multi-modal large model. It needs to be explained that the multi-modal large language model is a new research hotspot in recent years, which uses a powerful large language model as a brain to perform multi-modal tasks. Since the multi-modal large model adopted in the embodiments of the present disclosure is similar to the large model training method in the prior art, the training method of the multi-modal large model will not be described here.
[0067] For example, in a train fault detection system, there is a fault type called "inter-hook difference over-limit", which specifically refers to the abnormality of the inter-hook difference component, and only the inter-hook difference component can have such a fault. If a new target image to be detected is received, it will first be processed according to the simple fault type, that is, processed according to the same processing mode of the first type, because many faults can be quickly judged by direct feature recognition. However, if the component detection model identifies the inter-hook difference component in the image, the system will take additional steps, that is, further analysis through the multi-modal large model branch of complex faults. This branch uses advanced multi-modal data processing technology, combines image content and prior knowledge of components, and comprehensively judges whether the inter-hook difference component exists in the fault condition of over-limit. Such a double-checking mechanism ensures that even complex and difficult-to-observe faults can be accurately detected and identified, thereby improving the reliability and safety of train fault detection.
[0068] In this way, for the to-be-detected target fault of the attribute type of the second type, the efficient multi-modal large model is used for auxiliary judgment, the complex fault discrimination is realized without consuming too many computing resources, and then the fault detection efficiency can be improved. In this way, not only the accuracy and efficiency of fault detection can be improved, but also the adaptability and processing depth for complex faults are enhanced, thereby providing more comprehensive protection for the safe operation of the train.
[0069] In some optional embodiments, the second target image inference result and the prompt word are determined according to the business logic contained in the to-be-detected target image, and the prompt word is generated, including:
[0070] The business logic contained in the to-be-detected target image is analyzed.
[0071] The image inference result to be used is determined according to the analyzed business logic, and the second target image inference result is obtained.
[0072] The prompt word is generated according to the analyzed business logic using the prompt word generation template.
[0073] Specifically, when the prompt word is generated, the prompt word generation template is obtained according to the prior discrimination standard of the fault, and then the prompt word is generated using the prompt word generation template (which can also be referred to as Prompt).
[0074] In this way, the second target image inference result and the prompt word are determined according to the analyzed business logic, which can realize complex fault discrimination without consuming too many computing resources, thereby facilitating the judgment of complex faults and improving the fault detection efficiency.
[0075] In order to facilitate the scheme of the embodiments of the present application, the following examples are given, see Figure 3 , Figure 3 is a fault detection flowchart of the attribute type of the to-be-detected target image being the second type in an embodiment of the present disclosure. Assuming that it is judged whether there is a mutual hook difference overrun fault in the to-be-detected target image, after analyzing the to-be-detected target image, it is found that the business logic contained in the to-be-detected target image is that the mutual hook difference overrun fault occurs when the connection position of the car hook before and after the mutual hook difference station is shifted up and down, that is, the fault occurs. Then there can be as follows Figure 3The illustrated multimodal large model implicit logical discrimination process. First, through the parsed business logic, it can be known that the image inference result corresponding to the component positioning detection model needs to be used, and the image inference result is taken as the second target image inference result. Then, according to the business discrimination logic, the prompt word corresponding to the fault is generated, and through the above business logic, the prompt word is obtained: if the mutual hook difference hook connection position of the front and rear carriages occurs up and down offset, there is a "mutual hook difference overrun" fault, and whether the mutual hook difference in the figure exists the fault. Then, the prompt word and the second target image inference result are input into the multimodal large model, and the fault judgment result can be obtained through the multimodal large model, that is, the answer output by the multimodal large model is "yes, there is a mutual hook difference overrun fault".
[0076] In some optional embodiments, before obtaining the images taken by each detection station in the train and obtaining the corresponding image inference result according to the taken images, the method further comprises:
[0077] Obtaining a target image to be detected;
[0078] If the specified component structure object is not included in the target image to be detected, the attribute type of the target image to be detected is the first type;
[0079] If the specified component structure object is included in the target image to be detected, the attribute type of the target image to be detected is the second type.
[0080] Specifically, in the train detection process, the embodiments of the present disclosure first obtain images taken by each detection station, and then analyze these images to determine the attribute type of the target image to be detected. Specifically, it is first checked whether the target image to be detected includes a specific component structure image: if the specified component structure is not included, the attribute type of the image is considered to be the first type, i.e., a simple type fault, which usually has obvious fault characteristics and is easy to identify through preset rules; if the image includes the specified component structure, it is classified as the second type, i.e., a complex type fault, which may require more advanced analysis techniques for identification. Through this classification method, the system can apply appropriate detection strategies for different types of fault images, thereby improving the accuracy and efficiency of fault detection.
[0081] For example, in a train fault detection system, there is a fault type called "mutual hook difference overrun", which specifically refers to the abnormality of the mutual hook difference component, and only the mutual hook difference component can have this type of fault. If a new target image to be detected is received that includes the mutual hook difference component, the attribute type of the target image to be detected is classified as the second type, and for images of this type, simple fault type processing is performed first, and then in-depth analysis is performed through the multimodal large model branch of complex fault.
[0082] In this way, the images photographed by each detection station in the train are acquired, and the type of the to-be-detected target image is classified before the corresponding image inference result is obtained according to the photographed image, thereby facilitating rapid determination of the fault judgment mode to be adopted.
[0083] In order to facilitate overall understanding of the technical solutions of the present disclosure, see Figure 4 , Figure 4 is a flowchart of a fault detection method for a to-be-detected target image in an embodiment of the present disclosure. In this method, the images photographed by each detection station in the train are acquired, and the corresponding image inference result is obtained according to the photographed image. The image inference result includes the prediction result corresponding to the photographed image. The image inference result includes the 1-station image inference result, the 2-nd station image inference result, and the N-th station image inference result. Figure 2
[0084] If the attribute type of the to-be-detected target image is the first type, the business logic contained in the to-be-detected target image is analyzed to obtain a business analysis result; the image inference result to be used in the image inference result is determined according to the business analysis result, and a first target image inference result is obtained. The image inference result contained in the first target image inference result is subjected to name matching, threshold filtering, NMS, and IOU (intersection over union) calculation operations, respectively, to obtain respective processed image inference results. Then, the respective processed image inference results are subjected to logical operation processing according to actual needs, wherein the logical operation includes any one of logical AND, logical OR, and logical NOT, so that a detection result corresponding to the to-be-detected target image can be obtained.
[0085] If the type of the to-be-detected target image is the second type, after being processed by the processing mode of the first type, the to-be-detected target image is analyzed to obtain the business logic contained in the to-be-detected target image. Through the analyzed business logic, it can be known that the image inference result to be used, which is taken as a second target image inference result. Then, a prompt word corresponding to the fault is generated according to the business discrimination logic. Through the above business logic, the prompt word can be obtained, for example: whether component X in the image exists a fault according to the standard of a certain fault. Then, the prompt word and the second target image inference result are input into a multi-modal large model, and a fault judgment result can be obtained through the multi-modal large model.
[0086] The embodiment of the present disclosure starts from the characteristics and application scenarios of railway scene faults, and in view of the fact that the fault types in the real train scene are complex and interdependent, a multi-modal large model joint prediction method for train fault detection is proposed, which realizes the detection capability of cross-image and cross-camera faults and the complex fault implicit logic discrimination capability.
[0087] The device embodiment of the present application is introduced below, which can be used to execute the train fault detection method in the above-mentioned embodiments of the present application. For details not disclosed in the device embodiment of the present application, please refer to the above-mentioned embodiments of the train fault detection method of the present application.
[0088] The present disclosure also provides a train fault detection device 500, as shown, comprising: Figure 5
[0089] The acquisition module 501 is configured to acquire images captured by each detection station in the train, and obtain corresponding image inference results according to the captured images, wherein the image inference results contain prediction results corresponding to the captured images.
[0090] The first processing module 502 is configured to, if the attribute type of the target image to be detected is the first type, process the image inference results according to a preset rule to obtain a detection result corresponding to the target image to be detected.
[0091] The second processing module 503 is configured to, if the attribute type of the target image to be detected is the second type, simultaneously process the image inference results by using a multi-modal large model to obtain a fault result corresponding to the target image to be detected.
[0092] In some optional embodiments, the acquisition module 501 acquires images captured by each detection station in the train, and obtains corresponding image inference results according to the captured images, comprising:
[0093] Acquiring images captured by each detection station in the train.
[0094] Inputting the images captured by each detection station into different target detection models respectively, and obtaining image inference results corresponding to the images captured by each detection station by the target detection models.
[0095] In some optional embodiments, if the attribute type of the target image to be detected is the first type, the first processing module 502 processes the image inference results according to a preset rule to obtain a detection result corresponding to the target image to be detected, comprising:
[0096] If the type of the target image to be detected is the first type, analyzing the business logic contained in the target image to be detected to obtain a business analysis result.
[0097] Determining the image inference result to be used in the image inference result according to the business analysis result to obtain a first target image inference result.
[0098] Processing the first target image inference result according to the preset rule to obtain a detection result corresponding to the target image to be detected.
[0099] In some optional embodiments, the first processing module 502 processes the first target image inference result according to a preset rule to obtain a detection result corresponding to the target image to be detected, including:
[0100] If the first target image inference result includes at least two image inference results, each image inference result in the first target image inference result is processed respectively by using a preset rule to obtain a processed image inference result;
[0101] The processed image inference result is subjected to logical operation to obtain a detection result corresponding to the target image to be detected.
[0102] In some optional embodiments, if the attribute type of the target image to be detected is a second type, the second processing module 503 simultaneously processes the image inference result by using a multi-modal large model to obtain a fault result corresponding to the target image to be detected, including:
[0103] If the attribute type of the target image to be detected is the second type, after using the same processing mode as the first type, a second target image inference result is determined and a prompt word is generated according to the business logic contained in the target image to be detected;
[0104] The prompt word and the second target image inference result are input into the multi-modal large model, and a fault result corresponding to the target image to be detected is obtained through the multi-modal large model.
[0105] In some optional embodiments, the second processing module 503 determines the second target image inference result and generates the prompt word according to the business logic contained in the target image to be detected, including:
[0106] The business logic contained in the target image to be detected is parsed;
[0107] The image inference result to be used is determined according to the parsed business logic to obtain the second target image inference result;
[0108] The prompt word is generated by using a prompt word generation template according to the parsed business logic.
[0109] In some optional embodiments, before the acquisition module 501 acquires the images photographed by each detection station in the train and obtains the corresponding image inference result according to the photographed images, the acquisition module 501 is further used to acquire a target image to be detected;
[0110] If the target image to be detected does not contain a specified component structure object, the attribute type of the target image to be detected is the first type;
[0111] If the target image to be detected contains a specified component structure object, the attribute type of the target image to be detected is the second type.
[0112] In the technical solutions of the present disclosure, the acquisition, storage and application of user personal information are in line with relevant laws and regulations and do not violate public order and good customs.
[0113] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium and a computer program product.
[0114] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit implementations of the present disclosure described and / or claimed in this document.
[0115] As shown in Figure 6 The electronic device 600 includes a computing unit 601 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 602 or a computer program loaded into a random access memory (RAM) 603 from a storage unit 608. Various programs and data required for the operation of the device 600 can also be stored in the RAM 603. The computing unit 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0116] Various components in the device 600 are connected to the I / O interface 605, including an input unit 606, such as a keyboard, a mouse, etc., an output unit 607, such as various types of displays, speakers, etc., a storage unit 608, such as a magnetic disk, an optical disk, etc., and a communication unit 609, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 allows the device 600 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.
[0117] The computing unit 601 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 601 performs various methods and processes described above, such as the train fault detection method. For example, in some embodiments, the train fault detection method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded onto the RAM 603 and executed by the computing unit 601, one or more steps of the program distribution described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the train fault detection method by any other appropriate means, such as by means of firmware.
[0118] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0119] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, partially on a machine and partially on a remote machine or entirely on a remote machine or server.
[0120] In the context of this disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0121] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0122] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0123] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0124] It should be understood that the various forms of flow shown above can be used to reorder, add, or delete steps. For example, the steps described in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure can be achieved, which is not limited herein.
[0125] The above detailed description does not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A train fault detection method, comprising: obtaining images taken by each detection station in a train, and obtaining corresponding image inference results according to the taken images, wherein the image inference results include prediction results corresponding to the taken images; if an attribute type of a target image to be detected is a first type, processing the image inference results according to a preset rule to obtain a fault detection result corresponding to the target image to be detected, wherein if the target image to be detected does not include a specified component structure object, the attribute type of the target image to be detected is the first type; if the attribute type of the target image to be detected is a second type, after using the same processing mode as the first type, determining a second target image inference result and generating a prompt word according to a business logic included in the target image to be detected; inputting the prompt word and the second target image inference result into a multi-modal large model to obtain a fault detection result corresponding to the target image to be detected through the multi-modal large model, wherein if the target image to be detected includes the specified component structure object, the attribute type of the target image to be detected is the second type.
2. The method of claim 1, wherein, The obtaining images taken by each detection station in a train, and obtaining corresponding image inference results according to the taken images, comprises: obtaining images taken by each detection station in a train; inputting the images taken by each detection station into different target detection models respectively to obtain image inference results corresponding to the images taken by each detection station through the target detection models.
3. The method of claim 1, wherein, If the attribute type of the target image to be detected is the first type, processing the image inference results according to the preset rule to obtain the fault detection result corresponding to the target image to be detected, comprises: if the type of the target image to be detected is the first type, analyzing a business logic included in the target image to be detected to obtain a business analysis result; determining an image inference result to be used in the image inference results according to the business analysis result to obtain a first target image inference result; processing the first target image inference result according to the preset rule to obtain the fault detection result corresponding to the target image to be detected.
4. The method of claim 3, wherein, The processing the first target image inference result according to the preset rule to obtain the fault detection result corresponding to the target image to be detected, comprises: if the first target image inference result includes at least two image inference results, processing each image inference result in the first target image inference result according to the preset rule respectively to obtain processed image inference results; performing logical operation on the processed image inference results to obtain the fault detection result corresponding to the target image to be detected.
5. The method of claim 1, wherein, The determining a second target image inference result according to the business logic included in the target image to be detected and generating a prompt word, comprises: analyzing the business logic included in the target image to be detected; determining an image inference result to be used according to the analyzed business logic to obtain a second target image inference result; generating the prompt word according to the analyzed business logic using a prompt word generation template.
6. A train fault detection device, comprising: an acquisition module configured to acquire images captured by each detection station in a train and obtain corresponding image inference results according to the captured images, wherein the image inference results include prediction results corresponding to the captured images; a first processing module configured to, if an attribute type of a target image to be detected is a first type, process the image inference results according to a preset rule to obtain a fault detection result corresponding to the target image to be detected, wherein if the target image to be detected does not include a specified component structure object, the attribute type of the target image to be detected is the first type; a second processing module configured to, if the attribute type of the target image to be detected is a second type, after using the same processing mode as the first type, determine a second target image inference result and generate a prompt word according to a business logic included in the target image to be detected, and input the prompt word and the second target image inference result into a multi-modal large model to obtain a fault detection result corresponding to the target image to be detected by the multi-modal large model, wherein if the target image to be detected includes the specified component structure object, the attribute type of the target image to be detected is the second type.
7. The apparatus of claim 6, wherein, The acquisition module acquires images captured by each detection station in a train and obtains corresponding image inference results according to the captured images, comprising: acquiring images captured by each detection station in a train; inputting the images captured by each detection station into different target detection models respectively, and obtaining image inference results corresponding to the images captured by each detection station by the target detection models.
8. The apparatus of claim 6, wherein, The first processing module processes the image inference results according to a preset rule to obtain a fault detection result corresponding to the target image to be detected if the attribute type of the target image to be detected is a first type, comprising: if the type of the target image to be detected is the first type, analyzing a business logic included in the target image to be detected to obtain a business analysis result; determining an image inference result to be used in the image inference results according to the business analysis result to obtain a first target image inference result; processing the first target image inference result according to the preset rule to obtain a fault detection result corresponding to the target image to be detected.
9. The apparatus of claim 8, wherein, The first processing module processes the first target image inference result according to the preset rule to obtain a fault detection result corresponding to the target image to be detected, comprising: if the first target image inference result includes at least two image inference results, processing each image inference result in the first target image inference result according to the preset rule respectively to obtain processed image inference results; performing logical operations on the processed image inference results to obtain a fault detection result corresponding to the target image to be detected.
10. The apparatus of claim 6, wherein, The second processing module determines a second target image inference result and generates a prompt word according to a business logic included in the target image to be detected, comprising: analyzing the business logic included in the target image to be detected; According to the parsed business logic, determine the image inference result to be used, and obtain a second target image inference result; According to the parsed business logic, generate a prompt word using a prompt word generation template. 11.An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-5.
12. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method according to any one of claims 1-5. 13.A computer program product comprising a computer program which, when executed by a processor, implements the method according to any one of claims 1-5.
Citation Information
Patent Citations
Method and device for recognizing fault type
CN109245910A
Bolt looseness detection method, device and equipment and storage medium
CN113034456A