An adaptive multi-modal image recognition method and system based on artificial intelligence
By using an AI-based adaptive multimodal image recognition method, image information is acquired in real time and the detection type is predicted based on specification data. The appropriate image modality is selected for detection, which solves the problems of missed detection and false detection in complex production processes in intelligent manufacturing and achieves efficient and accurate quality control.
Patent Information
- Application Number
- CN202411933650.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2044-12-26
AI Technical Summary
Existing intelligent manufacturing inspection systems are unable to adapt to the dynamic changes in complex production processes, especially on production lines with multiple models or specifications, leading to missed or incorrect inspections and an inability to detect and handle quality problems in a timely manner.
An AI-based adaptive multimodal image recognition method is adopted to acquire image information of the target object in real time, predict the detection type based on specification data, select the appropriate image modality for detection, extract target features, and generate defect information that is sent to the management terminal in real time.
It improves the accuracy and efficiency of inspection, reduces the consumption of computing resources and hardware costs, ensures that quality inspection of each process is not missed, and maintains consistency and accuracy in large-scale production.
Smart Images

Figure CN119848457B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent manufacturing, in particular to a self-adaptive multi-modal image recognition method and system based on artificial intelligence. BACKGROUND
[0002] In the intelligent manufacturing industry, fixed detection processes and detection positions are adopted. The product of each process is sent to a fixed position for inspection when it passes through the process, and the inspection content is usually completed by manual or automatic detection equipment at a certain specific link. These detection systems usually complete the corresponding quality detection at the tail end of each process, and the position and state of the product at each link are relatively static, and the detection equipment is only responsible for the quality control of a specific link in the entire production process.
[0003] In the prior art, the quality detection of each process is limited to a specific position and stage. Once a product does not pass through a certain process, or there are different processing methods between processes, the system cannot effectively trace or detect the quality of the product. For complex production processes, especially production lines involving multiple product models or specifications, static detection methods at a single process position may result in missed or incomplete control of multiple process qualities. For example, if a process at a certain link is adjusted, or a certain product needs to skip certain processes due to quality problems, the existing detection system may not be able to adapt to these changes in time, resulting in missed detection or false detection. When a production line processes multiple models or specifications of products at the same time, the detection requirements of each product may be different, and the existing detection system cannot flexibly adapt to different detection standards between multiple processes. If a certain link detection is skipped or delayed, the system usually cannot automatically find the problem and take immediate remedial measures, resulting in missed quality control.
[0004] Therefore, the prior art has defects and needs to be improved. SUMMARY
[0005] In order to solve one or several problems in the prior art, the main purpose of the present application is to provide a self-adaptive multi-modal image recognition method and system based on artificial intelligence.
[0006] In order to achieve the above-mentioned purpose of the application, the present application provides a self-adaptive multi-modal image recognition method based on artificial intelligence, which comprises:
[0007] real-time acquisition of image information of a target detection object, and determination of whether the target detection object meets a detection condition according to the image information;
[0008] If the target detection object meets the detection condition, the specification data of the target detection object is acquired.
[0009] predict a detection type of the target detection object according to the specification data;
[0010] identify and generate a modality image required by the detection type based on the detection type, and extract a target feature of the modality image;
[0011] determine whether the target detection object meets a defect condition according to the target feature;
[0012] when the target detection object meets the defect condition, generate detection defect information and the modality image of the target detection object and send them to a management end.
[0013] The embodiment of the application further provides a self-adaptive multi-modality image recognition system based on artificial intelligence, comprising:
[0014] a first acquisition module, configured to acquire image information of a target detection object in real time, and determine whether the target detection object meets a detection condition according to the image information;
[0015] a second acquisition module, configured to acquire specification data of the target detection object if the target detection object meets the detection condition;
[0016] a prediction module, configured to predict a detection type of the target detection object according to the specification data;
[0017] an identification module, configured to identify and generate a modality image required by the detection type based on the detection type, and extract a target feature of the modality image;
[0018] a determination module, configured to determine whether the target detection object meets a defect condition according to the target feature;
[0019] a generation module, configured to generate detection defect information and the modality image of the target detection object and send them to a management end when the target detection object meets the defect condition.
[0020] The application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements steps of the method according to any one of the above embodiments when executing the computer program.
[0021] The application further provides a computer readable storage medium, which stores a computer program, and the computer program implements steps of the method according to any one of the above embodiments when executed by a processor.
[0022] The adaptive multi-modal image recognition method and system based on artificial intelligence provided by the embodiments of the present application can quickly screen out targets meeting the detection conditions based on image content, avoiding the processing of invalid data. The system intelligently predicts the detection type according to the specification data of the target detection object and selects the most suitable modal image for each detection type. For example, some objects may need high-resolution images to detect surface flaws, while others may need thermal imaging or X-rays to identify internal defects. Through this adaptive detection scheme, different types of defects can be more accurately identified, avoiding missed detection or false detection caused by using a uniform detection mode. By extracting the key features of the target and making predictions based on historical data and models, it is determined whether the target meets the defect standard. Compared with traditional manual inspection methods, this intelligent defect recognition not only improves the accuracy of detection, but also maintains consistency in large-scale production, reducing errors caused by human differences. Once a defect is found, the system can immediately generate detailed defect information and the corresponding modal image and send it to the management end in real time. By intelligently predicting the detection type and optimizing the image acquisition mode, the system avoids using a uniform image mode or excessive calculation for all objects, significantly reducing the consumption of computing resources and hardware costs. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 A flowchart of an adaptive multi-modal image recognition method based on artificial intelligence according to an embodiment of the present application;
[0024] Figure 2 A flowchart of an adaptive multi-modal image recognition method based on artificial intelligence according to an embodiment of the present application;
[0025] Figure 3 A structural schematic block diagram of an adaptive multi-modal image recognition system based on artificial intelligence according to an embodiment of the present application;
[0026] Figure 4 A structural schematic block diagram of a computer device according to an embodiment of the present application.
[0027] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solutions and advantages of the present application more clear, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and not to limit the present application.
[0029] Reference Figure 1 In the embodiments of the present application, an adaptive multi-modal image recognition method based on artificial intelligence is provided, which comprises:
[0030] S1, real-time acquisition of image information of a target detection object, and judging whether the target detection object meets a detection condition according to the image information;
[0031] S2, if the target detection object meets the detection condition, acquiring specification data of the target detection object;
[0032] S3, predicting a detection type of the target detection object according to the specification data;
[0033] S4, identifying and generating a modal image required by the detection type based on the detection type, and extracting a target feature of the modal image;
[0034] S5, judging whether the target detection object meets a condition of a defect according to the target feature;
[0035] S6, when the target detection object meets the condition of the defect, generating detection defect information and the modal image of the target detection object and sending them to a management end.
[0036] As described in steps S1-S3 above, during the production or inspection process, the image information of the target object (such as a component or finished product) is captured in real-time by image acquisition devices (such as cameras or sensors) installed at the production line or inspection points. These image information can be used for subsequent image analysis and processing, helping to determine whether the object meets the inspection conditions. By obtaining image information in real-time, the system can quickly respond to changes on the production line, reducing manual intervention and delays, and improving production efficiency. This also provides a first step of image data basis for subsequent image-based defect detection and specification analysis. Based on the preliminary analysis of image content, this step determines whether the target object meets the conditions for further detection. For example, certain objects may have limitations or requirements in terms of size, appearance, etc. These image features can be quickly analyzed by simple image processing algorithms (such as edge detection, morphological operations, etc.) and machine learning models to perform preliminary filtering or classification, ensuring that only target objects that meet certain conditions enter the next step of detailed detection process. This step can effectively filter out target objects that do not meet the requirements or do not need further detection, thereby improving detection efficiency and avoiding wasting computing resources and detection time. This is a preprocessing step that reduces unnecessary interference for subsequent more complex judgment and analysis. Once the target object passes the preliminary screening, the system will further obtain the specification data related to the target. Specification data can include information such as the size, material, production process of the object. It can be obtained through database query, interface with production ERP system or through pre-set standardized interface. Specification data plays a key role in predicting the type of detection and selecting the appropriate image modality. Obtaining the specification data of the target object makes the subsequent detection more intelligent and targeted. Different specifications of objects have different detection standards and focuses in different survival procedures (for example, some processing procedures pay more attention to dimensional accuracy, and others focus on surface defects), and through matching specification data for customized detection, the accuracy and efficiency of detection can be improved. According to the obtained specification data, the system can infer which type of detection is needed for the target object. For example, if the specification of the detection object is a certain high-precision component, the system may predict that its main detection focus is dimensional accuracy; if it is a large-scale production shell, it may focus on the detection of surface defects. Based on these predictions, the system will generate appropriate modal images (such as high-resolution, infrared images, etc.) for subsequent image recognition. This adaptive detection type prediction can greatly improve the flexibility and accuracy of detection. By targeting different types of detection, the use of fixed standardized methods to detect all objects is avoided, reducing unnecessary consumption of computing resources and improving the relevance and efficiency of detection.
[0037] As described in steps S4-S6 above, according to the detection type determined in the previous step, the system selects the most suitable image acquisition mode or technique (e.g., capturing details through a high-resolution camera, or using an infrared camera to capture thermal imaging information, or using X-rays for internal structure scanning). Different detection types require different image modalities to obtain target features, for example, the detection of surface defects may use high-definition visual images, while the detection of internal structures may rely on CT scans or X-ray imaging. This way of selecting appropriate image modalities based on detection types effectively avoids the "one-size-fits-all" solution, allowing each type of defect to be accurately identified in the most suitable way. By adopting the optimal image mode according to the different properties of the object (such as surface, internal, size, etc.), the accuracy of detection is improved. After obtaining the appropriate modal image, the system extracts key features in the image through image processing algorithms (such as edge detection, feature point extraction, deep learning models, etc.). These features include but are not limited to shape, texture, color distribution, size, defects, etc. Deep learning techniques (such as convolutional neural networks) can automatically extract more complex features from images to identify potential defects. Feature extraction is the core step of image recognition. Through in-depth analysis and learning, the most effective features related to defects can be automatically obtained from images, improving the intelligence and precision of detection. After feature extraction, the system uses pre-trained machine learning models (such as classifiers or regression models) to determine whether the target has defects. These models compare the target's features with historical data to assess whether the target meets the defect criteria. If the target features meet the definition of a defect, the system determines that the target object has a defect. By using machine learning models to automatically judge defects, the defect detection process is transformed from traditional manual judgment to an automated and precise process. This not only improves the efficiency of defect detection, but also maintains high consistency and accuracy in large-scale production. Once a defect is detected in the target, the system generates relevant defect information (such as defect type, location, severity, etc.) and corresponding modal images (specific images or schematics of the defect), and transmits this information to the management end or operators in real time. The management end can take appropriate measures (such as repair, rework, or scrap) based on this information. This step ensures rapid feedback and processing of defects, improving the controllability of product quality in the intelligent manufacturing process.
[0038] In a specific embodiment, in some production lines, although the product position of each process is fixed sometimes, due to the dynamic changes of the production line, the coordination between different processes, and the requirements of production efficiency, the state of the product between different processes may change at any time. For example, in an assembly line, a packaging line, or a multi-process alternating production process, the state of the product and the order of the process may not be completely fixed, and sometimes quality traceability and tracking are required. At this time, a single "static detection" may not be enough to meet the actual needs of the production line. Assuming that in a complex electronic product production line, multiple processes are involved: raw material preparation, component assembly, function test, appearance inspection, etc. Between these processes, the product may be moved to different positions or stages due to production scheduling, line adjustment, process order changes, and other factors. In this case, in order to ensure that the quality detection of each process is not missed, and at the same time the abnormal product can be traced, it is very meaningful to detect based on the dynamic state of the product at different positions. Through the intelligent detection system to automatically identify the state of the product at different positions, and automatically record the detection results at the end of each process, it can ensure that each link in the production process meets the quality requirements, and avoids missed detection or false matching. In actual production process, the product on the production line may change according to production demand or process adjustment. If a product is skipped or temporarily shelved in a link, the intelligent detection system can monitor the position and state of the product in real time to ensure that all products are subjected to necessary detection and avoid mistakes. In the case of a production line processing multiple products at the same time (such as mixed production of different models or specifications), the detection system can identify the product category in real time and adapt to different detection standards. The detection process of each process may be different, and through dynamic detection of the product at each position, it can ensure that each type of product is detected according to its required standard. If the production line fails or has quality problems, it can be accurately traced according to the detection data of each product at different process stages, locate the root cause of the problem, and quickly identify defects in the production process.
[0039] As mentioned above, the image content-based quick screening of targets meeting the detection conditions avoids the processing of invalid data. The system intelligently predicts the detection type based on the specification data of the target detection object and selects the most suitable modal image for each detection type. For example, some objects may require high-resolution images to detect surface flaws, while others may require thermal imaging or X-rays to identify internal defects. Through this adaptive detection scheme, different types of defects can be more accurately identified, avoiding missed or false detections due to the use of a uniform detection mode. By extracting the key features of the target and making predictions based on historical data and models, it is determined whether the target meets the defect standard. Compared with traditional manual inspection methods, this intelligent defect identification not only improves the accuracy of detection, but also maintains consistency in large-scale production and reduces errors caused by human differences. Once a defect is found, the system can immediately generate detailed defect information and the corresponding modal image and send it to the management end in real time. By intelligently predicting the detection type and optimizing the image acquisition mode, the system avoids using a uniform image mode or excessive computation for all objects, significantly reducing the consumption of computing resources and hardware costs.
[0040] Referring to Figure 2 In one embodiment, the method of extracting the target features of the modal image comprises:
[0041] S41, when the detection type is the size of the target detection object, marking the target detection object in the modal image and constructing a bounding box, the bounding box being used as a reference point for the size parameter;
[0042] S42, calculating a proportion parameter of the pixel size to the actual size of the bounding box according to the width and height of the bounding box;
[0043] S43, converting the pixel size of the target detection object to the actual physical size according to the proportion parameter;
[0044] S44, extracting the size feature based on the actual physical size of the target detection object;
[0045] S45, obtaining the target features based on the size feature.
[0046] As described in the previous steps, the target object is identified and located from the image. In this method, when the target size is detected, the target detection object is located and labeled by image processing algorithms (such as convolutional neural network CNN), marking the position of the object. The bounding box is a rectangular frame used to locate the object in the image, and its boundary is the external edge of the object. In this step, the target detection object will be framed and the spatial position of the target will be defined by the bounding box. This operation provides a reference point for subsequent size calculation. By labeling the target detection object and constructing the bounding box, the position of the target object in the image can be clearly obtained, providing accurate reference for subsequent size calculation. This step is the basis of size calculation and physical size conversion, ensuring that subsequent processing can focus on the target object and avoid interference from irrelevant objects. In the image, the width and height of the object are represented in pixels, but these pixel sizes cannot directly represent the actual physical size of the object. Therefore, first, the pixel size (width and height) of the target object in the image needs to be known, and then the ratio between the actual physical size and the pixel size needs to be calculated. The calculation method of the ratio parameter is to compare the actual size of the known object (for example, the standard reference size of the object) with the pixel size measured in the image to obtain a conversion coefficient. This ratio parameter represents the representative length of each pixel in the actual physical world. The calculation of the ratio parameter between the pixel size and the actual size ensures that the size information in the image can be converted into the real physical size in the real world. This ratio can be used for the size measurement of any other uncalibrated object, thereby realizing the accurate calculation of the size of the target object. Once the ratio parameter between the pixel size and the actual size is calculated, the next operation is to convert the pixel size (width and height) of the target object in the image by the ratio. The accurate physical size can be used for quality control, anomaly detection, and other tasks on the production line. Once the actual physical size of the target object is obtained, the relevant size features can be extracted. These features may include the aspect ratio, area, volume (for three-dimensional objects), standard deviation of size (if the variability of the object size needs to be considered), etc. These features help to describe the shape and size distribution characteristics of the object. The extraction of size features is crucial for subsequent defect identification. For example, if the size of a target exceeds the expected range, it may mean that the object has production defects or problems in the processing process. The extraction of size features not only helps to understand the compliance of the object, but also provides useful feature input for machine learning models. By extracting size features, detailed information can be provided for subsequent intelligent analysis, helping to identify objects or defects that do not meet specifications. By extracting target features, the system can comprehensively understand the properties of the object, not just relying on size features, but also combining other modal information (such as texture, color, etc.) to form a multi-modal target feature set.This is very helpful for target classification, defect detection and other tasks, and can greatly improve the robustness and accuracy of detection.
[0047] In an embodiment, the target feature in the image information is extracted, and the method further comprises:
[0048] When the detection type is the texture of the target detection object, marking the target detection object in the modal image;
[0049] Converting the target detection object into a gray scale image;
[0050] According to the result of processing, a gray scale co-occurrence matrix is constructed for adjacent pixel pairs in the horizontal direction of the gray scale image, for counting the occurrence frequency of each pair of gray scale values;
[0051] Based on the constructed gray scale co-occurrence matrix, the texture features of the target detection object are analyzed, including contrast, homogeneity and energy;
[0052] According to the analysis result, the distribution data of the gray scale values in the gray scale co-occurrence matrix are further analyzed for quantifying the gray scale distribution of the target detection object;
[0053] According to the distribution data of the gray scale values, the distribution area beyond the preset gray scale range is extracted as the target feature.
[0054] As mentioned above, converting the image to grayscale is a step in processing the texture features of the image. Grayscale removes color information and only retains brightness information. The grayscale value represents the lightness of each pixel in the image, using a value between 0 and 255 (8-bit grayscale). By removing color information, the complexity of the image is reduced, and texture analysis becomes more focused on the brightness changes of the object surface, without being disturbed by color. The process of converting to grayscale is done by taking a weighted average of the RGB values of each pixel, converting it to a single brightness value. Grayscale can more effectively reveal the texture features of the object surface. The conversion of grayscale makes subsequent texture feature calculation more simplified and focused, removing the complexity of color and strengthening the texture information of the image. The algorithm can focus on the structure and surface features in the image without being affected by color. The Gray-Level Co-occurrence Matrix (GLCM) is used to statistically analyze the spatial relationship between the gray values of pixels in the image. It captures the local texture features of the image by calculating the frequency of gray value pairs of pixel pairs (usually adjacent pixels) in a specific direction (such as horizontal, vertical or diagonal). In this method, after the target detection object is converted to a grayscale image, the algorithm constructs a gray-level co-occurrence matrix based on adjacent pixel pairs in the horizontal direction. This step provides useful data for further texture analysis by statistically analyzing the gray value combination of each adjacent pixel pair. Building a gray-level co-occurrence matrix can effectively capture the spatial information and structural features of the texture in the image. For example, in a two-dimensional plane, the smoothness, roughness, contrast and other features of the texture can be reflected by analyzing the gray-level co-occurrence matrix. This matrix is the core of texture feature extraction, providing key statistical basis for subsequent calculations such as contrast, homogeneity, energy and other features. Contrast reflects the local variation of pixel gray values in the image, indicating the roughness of the texture. A larger contrast value usually means that the texture is more complex and rough, and vice versa, indicating a smooth texture. Homogeneity describes the uniformity of the image texture. It reflects whether the distribution of gray values is uniform, and the larger the value, the smoother and more uniform the texture, and the smaller the value, the greater the non-uniformity in the texture. Energy reflects the regularity or consistency of the texture, which is the sum of the squares of the elements in the co-occurrence matrix. The larger the value, the more regular and consistent the texture. These texture features are usually calculated through the statistics of the gray-level co-occurrence matrix, and can effectively describe the texture, smoothness and complexity of the object surface. Analyzing these texture features helps to quantitatively describe the surface characteristics of the target object. In industrial detection, these texture features can help identify whether there are defects or irregularities on the surface of the object. For example, high contrast may indicate a rough surface, while low energy may indicate a lack of regularity on the surface. Through these features, the system can perform surface quality assessment, defect detection, etc.The distribution data of gray values reflects the frequency distribution of each gray level in the image. By analyzing the distribution of gray values, the algorithm can identify the gray areas or abnormal areas that appear more frequently on the surface of the object. These areas may represent certain specific areas of the object, such as certain defects or special texture patterns. In this step, the histogram of gray values or other statistical features of the gray co-occurrence matrix are usually analyzed to help quantify the texture distribution of the object. Analyzing the gray distribution helps to find possible abnormal areas on the surface of the object, especially when the distribution of certain gray values exceeds the pre-set normal range. These abnormal areas may represent defects or other quality problems. By quantifying these gray distributions, the system can achieve more accurate defect detection and quality evaluation. Based on the analysis of the gray value distribution, a pre-set gray range is set, and the gray distribution areas that exceed this range are detected and extracted. These areas that exceed the range usually represent abnormal areas in the image, such as surface defects, texture abnormalities, etc. Through this operation, the algorithm can identify and extract the target features that may be problem areas. The extraction of these gray abnormal areas helps the system to quickly locate potential defects or non-compliant parts during quality detection.
[0055] In an embodiment, the method of predicting the detection type of the target detection object according to the specification data comprises:
[0056] Obtaining the specification data of the target detection object and the image information of the current target detection object;
[0057] Based on the image information of the target detection object, obtaining the historical production data of the same type of detection object, and analyzing the change characteristics of the historical production data at different production stages according to the time sequence;
[0058] Inputting the specification data and the historical production data into a pre-set prediction model, predicting the current production stage of the target detection object through the prediction model, and outputting the prediction result;
[0059] Determining the detection type of the target detection object according to the predicted production stage, wherein the detection type includes the size, color and texture of the target detection object.
[0060] As mentioned above, specification data includes the size, weight, material, and other physical attributes of the target object. These data are obtained because they can provide important basic information for the prediction process. For example, size may affect the appearance of the object, the accuracy requirements of detection, and material may affect the optical properties (such as reflectivity, transparency, etc.) during detection. Image information is visual feature data that can intuitively reflect the appearance of the target object (such as color, texture, surface defects, etc.). These image data can be used to assist in analyzing the appearance and morphology of the target object, thereby further affecting the prediction of the detection type. By combining specification data and image information, the physical characteristics and visual performance of the target object can be comprehensively understood, which helps to more accurately predict its production stage and detection type. Historical production data can include performance data and production status of the target object at different times and different production stages. These data can reflect the change patterns of the target object during the production process. For example, certain sizes or colors may deviate in a certain production stage, or the texture may change with different production processes. Time series analysis is a technique for analyzing the change patterns of data over time. By sorting historical production data by time, the change trend of data in different production stages can be analyzed to find regular changes (such as the increase or decrease of certain features in certain stages, or the deviation pattern under certain production conditions). This helps to provide a reference for predicting the current production stage of the target object. Using historical production data and time series analysis can capture common change trends in the production process of the target object, and then infer the likelihood of the current production stage. This method helps to identify potential problems in advance and improves the accuracy and efficiency of prediction. The prediction model can use machine learning or deep learning algorithms to train with existing specification data and historical production data to learn the relationship between input data and target production stage. These models usually include regression analysis, decision tree, neural network, etc. Through learning from historical data, the prediction model can predict the production stage of the target object according to new input data (specification data and historical data). During the training process, the model will adjust the parameters repeatedly to fit the historical data and find the most suitable prediction rule. As data continues to accumulate, the accuracy of the model's prediction will also improve. By inputting specification data and historical production data into the prediction model, automated and intelligent production stage prediction can be achieved. Compared with manual judgment, the prediction model can make faster and more accurate predictions, reducing manual intervention and improving the efficiency of the production line.In the production process, the changes in size, color and texture are often important basis for quality detection. For example, some color deviation may mean material problems, texture problems may affect the use function, size deviation may affect the fitting accuracy, etc. Therefore, accurate detection type definition can help the factory to carry out targeted quality control. According to the predicted production stage to determine the detection type, the detection system can be more targeted to ensure that the production results of each stage meet the quality standards. For example, a stage may pay more attention to size accuracy, while another stage may pay more attention to appearance defects (such as uneven color or texture defects). This phased detection method helps to improve the detection efficiency of the production line and reduce the probability of false detection and missed detection.
[0061] In an embodiment, the method further comprises predicting a detection type of the target detection object according to the environmental data, the step comprising:
[0062] Real-time acquisition of environmental data in the range of the target detection object, extraction of target environmental features of the environmental data according to the specification data of the target detection object;
[0063] Constructing an analysis model, inputting the target environmental features and image information of the target detection object into the analysis model, analyzing the influence parameters of the target environmental features on the target detection object through the analysis model, and outputting the analysis results;
[0064] Based on the output influence parameters, determine the detection type of the target detection object.
[0065] As mentioned above, environmental data includes temperature, humidity, air pressure, light, air quality, and other external conditions. Environmental factors have a significant impact on the production and detection of target detection objects, especially in precision manufacturing and high-demand product production processes. For example, temperature changes can affect the expansion or contraction of materials, humidity can affect the degree of material corrosion, and light conditions can affect the quality of image acquisition. By obtaining environmental data in real time, dynamic monitoring of the current production environment can be ensured. The real-time nature of this data can reflect fluctuations in environmental conditions during the production process, providing immediate data support for subsequent prediction and analysis. Real-time acquisition of environmental data enables the entire system to react to the actual conditions of the current production environment and make timely adjustments. For example, if the temperature is too high or the humidity is too large, the production process or detection strategy may need to be adjusted immediately to avoid the impact of the environment on the target detection object. The specification data of the target detection object (such as size, material, weight, etc.) and environmental data are closely related. Different specifications may have different sensitivities to environmental conditions. For example, some materials may be very sensitive to temperature changes, while others may remain stable under extreme conditions. Therefore, by combining the specification data of the target detection object to extract the target environmental features from the environmental data, it is possible to more accurately identify which environmental factors are most important in the production or use of the detection object. These features may include the impact of temperature on size, the impact of humidity on surface quality, the impact of light on visual detection, etc. The extracted features will help provide appropriate input for subsequent model analysis. Extracting target environmental features can achieve an organic combination of environment and product specifications, making analysis more specific and targeted. Through effective identification of environmental impact factors, unnecessary external interference can be reduced, and the accuracy of model prediction can be improved. The purpose of building an analysis model is to combine target environmental features with image information of the target detection object for in-depth analysis. Typically, such analysis models may use machine learning, deep learning, or statistical regression methods. These models can learn and identify how environmental changes affect the detection results of the target detection object. For example, a deep neural network can be used to process complex image information and extract complex relationships between environmental conditions and image features. Image information itself contains appearance features of the target detection object, such as color, texture, shape, etc., while environmental features (such as light intensity, reflectivity, etc.) directly affect the performance of these visual features. For example, strong light may cause image overexposure, while high temperature may cause changes to the surface of the material, affecting the clarity of the texture. By inputting both target environmental features and image information into the model, the model can fully consider the impact of environmental changes on the target detection object during analysis. This comprehensive analysis can improve the accuracy of prediction and reduce errors caused by external environmental interference. By inputting environmental features and image information into the analysis model, the system can gain a deep understanding of how environmental factors interact with the appearance features of the target detection object.This method can identify which environmental conditions have a significant impact on the target detection object, thereby achieving more accurate prediction and detection. The impact parameters are variables analyzed by the model based on environmental data and image information, which represent the influence of environmental factors on the specific properties of the target detection object. For example, temperature changes may affect the size of the target detection object, humidity changes may affect the surface gloss or reflectivity of the material, etc. The task of the analysis model is to extract these impact parameters from the input data and evaluate their actual impact on the target detection object. The analysis results output by the model include the degree and direction of the influence of each environmental factor on the target detection object (such as affecting the size to increase or decrease, the possibility of surface defects, etc.). These results help determine the type of detection and detection criteria for the next detection process to ensure the relevance and efficiency of the detection process. The output impact parameters can provide scientific basis for the determination of detection type. By understanding the influence of environmental factors on the target detection object, it can be accurately judged whether the detection strategy or inspection items need to be adjusted. For example, if the temperature is too high, the model may predict that the size of the target detection object is biased, and the detection process may focus on size accuracy; if the humidity is too large, more attention may be needed to surface defects or corrosion problems. Based on the output impact parameters, the most appropriate detection items can be determined. For example, during the production process, temperature changes have a greater impact on the size changes of the target detection object, so the detection may focus on size accuracy and shape deviation; if the lighting conditions are poor, image acquisition methods may need to be adjusted or texture defects may need to be focused on. By analyzing the impact parameters, the system can dynamically adjust the detection strategy to ensure that the target detection object can achieve the best detection effect under the current environmental conditions. Determining the detection type based on the impact parameters can improve the accuracy and efficiency of detection. The advantage of this is that the detection strategy is more flexible, not only considering the requirements of the production stage, but also adjusting in real time according to environmental conditions, minimizing the interference of external environment on the detection results.
[0066] In an embodiment, the determining whether the target detection object meets the detection condition according to the image information comprises:
[0067] extracting visible features of the target detection object in the image information;
[0068] determining whether abnormal features appear according to the visible features;
[0069] when the visible features appear abnormal features, it is determined that the target detection object does not meet the detection condition, and the detection defect information of the target detection object is sent to the management end based on the determination result;
[0070] when the visible features do not appear abnormal features, it is determined that the target detection object meets the detection condition.
[0071] As mentioned above, image information contains various visual features of the target detection object, such as shape, color, texture, edge, size, etc. These visual features are key factors in evaluating whether the target detection object meets the design standards. Through computer vision technology, these features can be extracted from the image and analyzed. "Visible features" refer to physical properties related to the target detection object that can be directly observed in the image. For example, the smoothness of the surface of the target detection object, the symmetry of the shape, the consistency of the color, surface defects (such as cracks, scratches), etc., are typical visible features. These features directly affect the quality and function of the detection object. Generally, image processing algorithms (such as edge detection, texture analysis, color distribution analysis, etc.) are used to extract these visible features by analyzing the pixel data of the image. Convolutional Neural Networks (CNN) in deep learning are also a common method for image feature extraction, which can automatically learn and extract representative features from images. Extracting visible features from images is the basis for determining whether the target detection object meets the detection conditions. By accurately extracting these features, subsequent analysis can more effectively identify quality problems of the target detection object. This method can convert image information into feature data that machines can understand and process, thereby performing accurate analysis and judgment. Once the visible features of the target detection object are extracted, the next step is to determine whether these features meet the pre-defined standards or range. If some features (such as shape defects, surface scratches, color unevenness, etc.) deviate from the normal standard, they will be considered "abnormal features". This judgment is usually made by comparing with known standard features, or using machine learning models to identify which features do not conform to normal conditions. In practical applications, the standard of abnormal features is usually based on historical data, product specifications, or regulations in the production process. For example, for a mechanical part, there may be a size tolerance requirement, and exceeding the tolerance range will be considered abnormal; for surface-coated products, scratches or bubbles are considered abnormal features. By quantifying and classifying image features, the actual extracted features are compared with pre-defined standards. This can be done through statistical methods, rule engines, or trained machine learning models. By judging the visible features, potential defects of the target detection object can be discovered in a timely manner. This is a key step in quality control, which ensures that only products that meet the standards can pass the detection and prevents unqualified products from entering the market. Abnormal feature judgment not only depends on the deviation of features, but also considers environmental and process factors, which can dynamically adapt to different production conditions and product specifications, reducing the error and subjectivity of manual inspection. If the abnormality of the image features is confirmed, it means that the target detection object has defects and cannot meet the design and production requirements. Therefore, it is judged that it "does not meet the detection conditions". This is usually judged according to pre-defined quality standards and product requirements. Once unqualified products are found, the system needs to feedback the detection results to the management end in a timely manner.This step is usually carried out through an automated system, and the management end can be a quality monitoring system, a production management system, etc., for further processing, recording and analyzing defects. Defect information may include defect type, location, severity, etc., and the system can automatically generate reports or notify relevant personnel for follow-up processing. Through automated data transmission and feedback, the management end can receive defect information in real time and take timely measures. This process reduces the time for manual intervention, improves response speed and processing efficiency. This step ensures real-time feedback and decision-making of detection results. Through automation, defect information is transmitted to the management end in a timely manner, not only improving detection efficiency, but also ensuring quick response to unqualified products during production, reducing the risk of production and quality management. This mechanism helps to identify and solve problems in advance, thereby avoiding potential quality accidents and ensuring product quality stability and qualification rate. If the visible features extracted from the image do not appear abnormal, it can be inferred that the target detection object meets the detection standard. This means that the target detection object's features such as appearance, size, shape, etc. are within the acceptable range and meet the design requirements and quality control requirements. In actual detection, if no abnormal features are found, the system can consider that the detected object has passed the detection condition judgment. At this time, it can enter the production circulation or the next process stage. This step of judgment is a "pass" signal, which means that the target detection object meets all necessary quality standards under the current environmental conditions. The system can automatically record this result and send a pass information to the relevant system.
[0072] In an embodiment, the target environmental features include: high-temperature environmental features, high-humidity environmental features, and high-vibration environmental features. High-temperature environment: focus on surface texture. High temperature usually causes material expansion, deformation, and surface may appear thermal expansion texture or cracks unrelated to the material itself. If the system can sense the environmental temperature and feedback in real time through temperature sensors, the detection system can then take "surface texture" as the main feature, especially those thermal textures, cracks, etc. that are prone to form under temperature changes. For example, by monitoring the temperature distribution through thermal imaging cameras or infrared sensors, the detection system can identify cracks or expansion caused by thermal stress and distinguish them from defects in the manufacturing process. High-humidity environment: focus on color features. In a high-humidity environment, materials (especially certain plastics, wood, or coatings) may absorb moisture, swell, or discolor, and the surface may appear color changes, discoloration, or other physical deformations caused by moisture. If the humidity sensor detects a high humidity level, the system can focus on color features (such as color difference, reflectivity, etc.), which may not be caused by manufacturing defects but by physical changes caused by humidity. By analyzing the changes in color, it can be avoided to consider surface color unevenness caused by humidity changes as defects. Vibration environment: focus on micro-deformation or cracks. Vibration can cause products to deform slightly or crack, especially in dynamic environments, where micro-defects in products may become more apparent or expand under the influence of vibration. By monitoring vibration data in real time, the detection system can focus on "deformation" or "micro-cracks" as target features for analysis, especially in environments with high vibration intensity, where these defects may become apparent under the influence of vibration, which conventional detection may miss.
[0073] Referring to Figure 3 In the embodiments of the present application, an adaptive multi-modal image recognition system based on artificial intelligence is also provided, comprising:
[0074] A first acquisition module 1 is configured to acquire image information of a target detection object in real time, and determine whether the target detection object meets detection conditions according to the image information.
[0075] A second acquisition module 2 is configured to acquire specification data of the target detection object if the target detection object meets the detection conditions.
[0076] A prediction module 3 is configured to predict a detection type of the target detection object according to the specification data.
[0077] An identification module 4 is configured to identify and generate a modal image required by the detection type based on the detection type, and extract a target feature of the modal image.
[0078] A judgment module 5 is configured to determine whether the target detection object meets the conditions of defects according to the target feature.
[0079] The generating module 6 is configured to generate the detection defect information and the modality image of the target detection object and send them to the management end when the target detection object meets the defect condition.
[0080] As described above, it can be understood that each component of the adaptive multi-modal image recognition system based on artificial intelligence provided in the present application can realize the function of any one of the adaptive multi-modal image recognition methods based on artificial intelligence as described above, and the specific structure will not be described again.
[0081] Referring to Figure 4 The present application also provides a computer device, which can be a server, and the internal structure thereof can be as shown in Figure 4 The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store monitoring data and other data. The network interface of the computer device is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement an adaptive multi-modal image recognition method based on artificial intelligence.
[0082] The processor executes the adaptive multi-modal image recognition method based on artificial intelligence as described above, including: acquiring image information of a target detection object in real time, determining whether the target detection object meets a detection condition according to the image information; if the target detection object meets the detection condition, acquiring specification data of the target detection object; predicting a detection type of the target detection object according to the specification data; identifying and generating a modality image required by the detection type based on the detection type, extracting a target feature of the modality image; determining whether the target detection object meets a defect condition according to the target feature; and generating detection defect information and the modality image of the target detection object and sending them to a management end when the target detection object meets the defect condition.
[0083] An embodiment of the present application also provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement an artificial intelligence based adaptive multi-modal image recognition method, including the steps of: acquiring image information of a target detection object in real time, determining whether the target detection object meets a detection condition according to the image information; if the target detection object meets the detection condition, acquiring specification data of the target detection object; predicting a detection type of the target detection object according to the specification data; identifying and generating a modal image required by the detection type based on the detection type, and extracting a target feature of the modal image; determining whether the target detection object meets a defect condition according to the target feature; and when the target detection object meets the defect condition, generating detection defect information and the modal image of the target detection object and sending them to a management end.
[0084] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, the processes of the above-mentioned embodiments can be included. Any reference to memory, storage, database or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM) and memory bus dynamic RAM (RDRAM) and the like.
[0085] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover non-exclusive inclusions, such that processes, devices, articles or methods including a series of elements not only include those elements, but also include other elements not explicitly listed, or inherent to such processes, devices, articles or methods. Without more limitations, the element defined by the statement "including a" does not exclude the presence of another identical element in the process, device, article or method including the element.
[0086] The above merely provides the preferred embodiments of the present application, and is not intended to limit the patent scope of the present application. Any equivalent structure or equivalent flowchart transformation based on the content of the present application specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. An adaptive multimodal image recognition method based on artificial intelligence, characterized in that, The method includes: Real-time acquisition of image information of the target object, and determination of whether the target object meets the detection conditions based on the image information; If the target analyte meets the detection conditions, obtain the specification data of the target analyte; Predict the detection type of the target analyte based on the specified data; Based on the detection type, identify and generate the modal image required for the detection type, and extract the target features of the modal image; Based on the target characteristics, determine whether the target object meets the defect conditions; When the target object meets the defect conditions, the defect information and modal image of the target object are generated and sent to the management terminal. The method for predicting the detection type of the target object based on the specification data includes: acquiring the specification data of the target object and the current image information of the target object; acquiring historical production data of the same type of target object based on the image information of the target object, and analyzing the variation characteristics of the historical production data at different production stages according to the time series; inputting the specification data and historical production data into a preset prediction model, predicting the current production stage of the target object through the prediction model, and outputting the prediction result; determining the detection type of the target object based on the predicted production stage, wherein the detection type includes the size, color, and texture of the target object; the historical production data includes the performance data and production status of the target object at different periods and different production stages, which can reflect the possible change patterns of the target object during the production process; Environmental data includes temperature, humidity, air pressure, light intensity, air quality, and external conditions. In precision manufacturing and high-requirement product production processes, temperature changes may affect material expansion or contraction, humidity may affect the degree of material corrosion, and light conditions may affect image acquisition quality. By acquiring environmental data in real time, dynamic monitoring of the current production environment can be achieved. This real-time data reflects fluctuations in environmental conditions during production. When the temperature or humidity is too high, the production process or detection strategy needs to be adjusted immediately to avoid the impact of the environment on the target object being detected. Specification data of the target object being detected includes dimensions, material, and weight. Specification data is closely related to environmental data; different specifications have different sensitivities to environmental conditions. Therefore, by combining the specification data of the target object being detected, target environmental features can be extracted from the environmental data. Extracting target environmental features achieves an organic combination of environment and product specifications.
2. The adaptive multimodal image recognition method based on artificial intelligence according to claim 1, characterized in that, The method for extracting target features from the modality image includes: When the detection type is the size of the target object, the target object in the modal image is marked and a bounding box is constructed, the bounding box being used as a reference point for the size parameter; Calculate the ratio of the pixel size of the bounding box to its actual size based on the width and height of the bounding box; According to the aforementioned proportional parameters, the pixel size of the target object is converted into its actual physical size; The size features are extracted based on the actual physical size of the target object. Based on the size characteristics, the target features are obtained.
3. The adaptive multimodal image recognition method based on artificial intelligence according to claim 2, characterized in that, The method for extracting target features from the image information further includes: When the detection type is the texture of the target object, the target object in the modal image is marked; The target object is converted into a grayscale image; Based on the processing results, a gray-level co-occurrence matrix is constructed for adjacent pixel pairs in the horizontal direction of the grayscale image to count the frequency of occurrence of each pair of grayscale values. Based on the constructed gray-level co-occurrence matrix, the texture features of the target object are analyzed, including contrast, homogeneity, and energy. Based on the analysis results, further analysis of the gray value distribution data in the gray co-occurrence matrix is used to quantify the gray distribution of the target detection object; Based on the distribution data of grayscale values, extract the distribution areas that exceed the preset grayscale range as target features.
4. The adaptive multimodal image recognition method based on artificial intelligence according to claim 1, characterized in that, The method for determining whether the target object meets the detection conditions based on image information includes: Extract the visible features of the target object from the image information; Determine whether abnormal features are present based on the visible features; When the visible features show abnormal features, it is determined that the target object does not meet the detection conditions, and the detection defect information of the target object is sent to the management terminal based on the judgment result; If no abnormal features are found in the visible features, the target object is determined to meet the detection conditions.
5. The adaptive multimodal image recognition method based on artificial intelligence according to claim 1, characterized in that, The target environmental characteristics include: high temperature environment characteristics, high humidity environment characteristics, and high vibration environment characteristics.
6. An artificial intelligence-based adaptive multimodal image recognition system, used in the method described in any one of claims 1-5, characterized in that, include: The first acquisition module is used to acquire image information of the target object in real time and determine whether the target object meets the detection conditions based on the image information. The second acquisition module is used to acquire the specification data of the target detection object if the target detection object meets the detection conditions; The prediction module is used to predict the detection type of the target analyte based on the specification data. The recognition module is used to recognize and generate a modal image required for the detection type based on the detection type, and extract the target features of the modal image; The judgment module is used to determine whether the target object being detected meets the conditions for a defect based on the target features. The generation module is used to generate detection defect information and modal images of the target object and send them to the management terminal when the target object meets the defect conditions.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Device defect detection method and device, computer equipment and storage medium
CN117575995A
Product defect detection method and device based on intelligent light source
CN118896958A
Data processing method and system based on industrial internet
CN119089351A