Industrial scene recognition method and device, electronic equipment and storage medium
By using industrial video streams and pre-trained scene recognition models in industrial scene recognition, combined with dynamic change detection and comparative image feature comparison methods, the problem of poor accuracy and real-time performance in industrial scene recognition is solved, and more efficient industrial scene recognition is achieved.
Patent Information
- Application Number
- CN202510167669.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-14
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has problems with poor monitoring and identification accuracy and real-time performance in industrial scenario recognition, and it is impossible to effectively identify complex and changeable industrial scenarios.
By obtaining the industrial video stream in the target industrial scenario, inputting it into the pre-trained scene recognition model, and multi-frame industrial images are recognized in chronological order. If there is a dynamic change in the industrial object in the industrial image at the current time, it is determined as the target industrial image, and feature extraction and comparison with the comparative industrial image at the adjacent time to generate scene description information.
Improve the accuracy and real-timeness of industrial scene recognition, and better understand the target industrial image, thereby describing industrial scenes more accurately.
Smart Images

Figure CN120147950A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of industrial image processing, and in particular, to an industrial scenario recognition method, apparatus, electronic device, and storage medium. Background Art
[0002] In different industrial scenarios, relevant personnel can carry out corresponding industrial production activities, and industrial equipment also operates in different production states. Based on this, industrial scenario recognition refers to the recognition of personnel operations, equipment changes, etc. involved in this process. In this way, all aspects of industrial production activities can be comprehensively understood and mastered, which is convenient for subsequent production optimization.
[0003] In the related art, the recognition of industrial scenarios depends on pre-set rules, so it can only monitor scenarios within the scope of the recognition rules. However, the actual situation is complex and changeable. In the process of monitoring and recognizing industrial scenarios in the related art, there are problems of poor monitoring and recognition accuracy and real-time performance. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose an industrial scenario recognition method, apparatus, electronic device, and storage medium, aiming to improve the accuracy and real-time performance of industrial scenario recognition.
[0005] To achieve the above object, a first aspect of the embodiments of the present application proposes an industrial scenario recognition method, including:
[0006] Obtain an industrial video stream in a target industrial scenario, where there are multiple industrial objects in the target industrial scenario, and the industrial video stream includes multiple industrial images collected from the target industrial scenario from a past moment to the current moment;
[0007] Input the industrial video stream into a pre-trained scenario recognition model, and recognize the multiple industrial images in chronological order;
[0008] If there are dynamic changes in the industrial objects in the industrial image at the current moment, determine the industrial image at the current moment as the target industrial image, and determine the comparison industrial image of the target industrial image at the adjacent moment;
[0009] Extract the first image feature of the target industrial image, and extract the second image feature of the comparison industrial image, and describe the target industrial image based on the difference between the first image feature and the second image feature, and output the scene description information of the target industrial scenario.
[0010] In some embodiments, the industrial image includes a first industrial image and a second industrial image;
[0011] If there are dynamic changes in the industrial objects in the industrial image at the current moment, determining the industrial image at the current moment as the target industrial image includes:
[0012] Selecting multiple first industrial images from the historical moments of the industrial video stream, where the image features of the multiple first industrial images are the same at the preset positions;
[0013] Performing inter-frame fusion processing on multiple first industrial image frames to obtain a first updated image, and determining first foreground information from the first updated image;
[0014] Selecting a second industrial image from the current moment of the industrial video stream, and determining second foreground information from the second industrial image;
[0015] If the first foreground information and the second foreground information do not match, determining that there are dynamic changes in the second industrial image, and determining the second industrial image at the current moment as the target industrial image.
[0016] In some embodiments, extracting first image features of the target industrial image, and extracting second image features of the comparison industrial image, and performing scene description on the target industrial image based on the differences between the first image features and the second image features, and outputting scene description information of the target industrial scene, including:
[0017] Performing feature encoding processing on the target industrial image to obtain a first encoded feature, and performing feature encoding processing on the comparison industrial image to obtain a second encoded feature;
[0018] Performing feature decoding processing on the first encoded feature to obtain first image features, and performing feature decoding processing on the second encoded feature to obtain second image features;
[0019] Generating comparison description information of the comparison industrial image based on the second image features, and determining the scene description information of the target industrial image based on the first image features and the comparison description information.
[0020] In some embodiments, performing feature encoding processing on the target industrial image to obtain a first encoded feature includes:
[0021] Performing feature encoding processing on the second foreground information of the target industrial image based on a preset first weight value to obtain a foreground first encoded feature;
[0022] Performing feature encoding processing on other image information of the target industrial image except the second foreground information based on a preset second weight value to obtain a background first encoded feature, where the first weight value is greater than the second weight value;
[0023] Obtaining the first encoded feature based on the foreground first encoded feature and the background first encoded feature.
[0024] In some embodiments, performing feature decoding processing on the first encoded feature to obtain a first decoded feature includes:
[0025] Performing first feature decoding processing on the first encoded feature based on a first long short-term neural network provided with a tangent activation function to obtain an initial decoded feature;
[0026] Performing second feature decoding processing on the initial encoded feature based on a second long short-term neural network provided with a tangent activation function to obtain the first decoded feature.
[0027] In some embodiments, obtaining a sample industrial video stream in a sample industrial scenario and corresponding sample label information of the sample industrial video stream, where there are multiple sample industrial objects in the sample industrial scenario, and the sample industrial video stream includes multiple frames of sample industrial images collected from the sample industrial scenario from a past moment to the current moment;
[0028] Inputting the sample industrial video stream and the sample label information into a pre-constructed initial scene recognition model to perform recognition on multiple frames of sample industrial images in chronological order;
[0029] If there are dynamic changes in the sample industrial objects in the sample industrial image at the current moment, determining the sample industrial image at the current moment as the sample industrial image and determining the sample comparison industrial image of the sample industrial image at the adjacent moment;
[0030] Extracting the sample first image feature of the sample industrial image and extracting the sample second image feature of the sample comparison industrial image, and describing the sample industrial image based on the difference between the sample first image feature and the sample second image feature, and outputting the sample scene description information of the sample industrial scenario;
[0031] Calculating a sample loss value representing the probability distribution difference between the sample scene description information and the sample label information, and calculating a sample evaluation value representing the description quality difference between the sample scene description information and the sample label information;
[0032] Adjusting the model parameters of the initial scene recognition model according to the sample loss value and the sample evaluation value until the training stop condition is reached, and obtaining a trained scene recognition model.
[0033] In some embodiments, calculating a sample loss value representing the probability distribution difference between the sample scene description information and the sample label information includes:
[0034] Performing mapping processing on the sample label information to obtain a sample label feature;
[0035] Calculating the probability distribution difference between the sample first decoded feature and the sample label feature based on a preset total loss function to obtain the sample loss value.
[0036] In some embodiments, mapping processing is performed on the sample label information to obtain sample label features, including:
[0037] Obtain the industrial vocabulary corresponding to the sample industrial video stream, where the industrial vocabulary includes multiple industrial-specific words;
[0038] Compare the sample label information with the industrial-specific words, and filter out redundant words from the sample label information to obtain the updated sample label information;
[0039] Perform mapping processing on the updated sample label information to obtain sample label features.
[0040] In some embodiments, the sample label information is determined through the following steps, and the steps include:
[0041] Based on a preset description template, describe different types of sample industrial objects in the sample industrial video stream to obtain the description information of each sample industrial object under the corresponding type, where the description templates corresponding to different types of sample industrial objects are different;
[0042] Merge the description information corresponding to each sample industrial object to obtain the sample label information.
[0043] In some embodiments, a sample evaluation value representing the description quality difference between the sample scene description information and the sample label information is calculated, including:
[0044] Parse the sample scene description information and the sample label information respectively to obtain a plurality of first words and a plurality of second words;
[0045] Based on a preset first evaluation function, calculate the sample evaluation value representing the matching degree between each first word and the corresponding second word;
[0046] Alternatively, based on a preset second evaluation function, calculate the sample evaluation value representing the coverage degree between all the first words and all the second words.
[0047] To achieve the above object, a second aspect of the embodiments of the present application proposes an industrial scene recognition device, including:
[0048] An acquisition module, configured to acquire an industrial video stream in a target industrial scene, where the target industrial scene contains multiple industrial objects, and the industrial video stream includes multiple industrial images collected from the target industrial scene from a past moment to the current moment;
[0049] An input module, configured to input the industrial video stream into a pre-trained scene recognition model, and perform recognition on multiple industrial images in chronological order;
[0050] An industrial image determination module, configured to determine the industrial image at the current moment as a target industrial image and determine a comparison industrial image of the target industrial image at an adjacent moment if there are dynamic changes in industrial objects in the industrial image at the current moment;
[0051] A target module, configured to extract first image features of the target industrial image, extract second image features of the comparison industrial image, describe the scene of the target industrial image based on the differences between the first image features and the second image features, and output scene description information of the target industrial scene.
[0052] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the industrial scene recognition method in the first aspect is implemented.
[0053] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the industrial scene recognition method in the first aspect is implemented.
[0054] The industrial scene recognition method, device, electronic device, and storage medium provided by the present application. The method obtains an industrial video stream in a target industrial scene, where there are multiple industrial objects in the target industrial scene, and the industrial video stream includes multiple industrial images collected from the target industrial scene from a past moment to the current moment; inputs the industrial video stream into a pre-trained scene recognition model to recognize the multiple industrial images in chronological order; if there are dynamic changes in industrial objects in the industrial image at the current moment, determines the industrial image at the current moment as a target industrial image and determines a comparison industrial image of the target industrial image at an adjacent moment; extracts first image features of the target industrial image, extracts second image features of the comparison industrial image, and describes the scene of the target industrial image based on the differences between the first image features and the second image features, and outputs scene description information of the target industrial scene. Since the second image features corresponding to the comparison industrial image usually represent the known static situation of the target industrial scene, on the basis of considering the first image features corresponding to the target industrial image, the second image features corresponding to the comparison industrial image are synchronously considered to help the scene recognition model better understand the target industrial image, thereby improving the accuracy and real-time performance of recognizing the target industrial scene. Description of the Drawings
[0055] Figure 1 It is a schematic diagram of an application scenario of the industrial scene recognition device provided by the embodiments of the present application;
[0056] Figure 2It is an optional flowchart of the industrial scenario recognition method provided by an embodiment of the present application;
[0057] Figure 3 It is an optional schematic diagram of an industrial image of the industrial scenario recognition method provided by an embodiment of the present application;
[0058] Figure 4 It is an optional schematic diagram of a target industrial image and its scenario description information of the industrial scenario recognition method provided by an embodiment of the present application;
[0059] Figure 5 It is an optional schematic diagram of the training of a scenario recognition model of the industrial scenario recognition method provided by an embodiment of the present application;
[0060] Figure 6 It is an optional flowchart of the industrial scenario recognition device provided by an embodiment of the present application;
[0061] Figure 7 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0063] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different module division in the device or a different order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0064] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0065] Under different industrial scenarios, relevant personnel can carry out corresponding industrial production activities, and industrial equipment also operates in different production states. Based on this, industrial scenario recognition refers to the recognition of personnel operations, equipment changes, etc. involved in this process. In this way, all aspects of industrial production activities can be comprehensively understood and mastered, which is convenient for subsequent production optimization.
[0066] In the related art, the recognition of industrial scenarios relies on pre-set rules, so only the scenarios within the recognition rules can be monitored. However, the actual situation is complex and changeable. In the process of monitoring and recognizing industrial scenarios, the related art has problems of poor accuracy and real-time performance in monitoring and recognition.
[0067] Based on this, the embodiments of the present application provide an industrial scenario recognition method, device, electronic device, and storage medium, aiming to improve the accuracy and real-time performance of industrial scenario recognition.
[0068] Exemplarily, as Figure 1 shown, Figure 1 is a schematic diagram of the application scenario of the industrial scenario recognition device provided by the embodiments of the present application. In an optional application scenario, the client 11 is communicatively connected to the server 12, and the industrial scenario recognition device proposed by the embodiments of the present application is deployed in the server 12. Moreover, a trained scenario recognition model is set in the industrial scenario recognition device; the user can input an industrial video stream to be processed through the client 11, and the scenario recognition model will recognize multiple frames of industrial images in chronological order; if there are dynamic changes in the industrial objects in the industrial image at the current moment, the industrial image at the current moment is determined as the target industrial image, and the comparison industrial image at the adjacent moment of the target industrial image is determined; then, the first image feature of the target industrial image is extracted, and the second image feature of the comparison industrial image is extracted, and the target industrial image is described based on the difference between the first image feature and the second image feature, and the scene description information of the target industrial scene is output.
[0069] It should be noted that in the embodiments of the present application, when it comes to information related to user characteristics such as user basic information or user identity, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained first. After clearly obtaining the user's separate permission or separate consent, the necessary data for the normal operation of the embodiments of the present application will be obtained. For example, before obtaining the industrial video stream to be processed, the consent of the relevant regulatory personnel to whom the industrial video stream belongs will be obtained first, otherwise the industrial video stream data that cannot be applied to the embodiments of the present application will be obtained. In addition, other relevant data obtained by the industrial scenario recognition device in the embodiments of the present application are all authorized data, which will not be elaborated here one by one.
[0070] In the embodiments of the present application, the description will be made from the dimension of the industrial scenario recognition device, and the industrial scenario recognition device can be integrated in a computer device, such as a server. As Figure 2 shown, Figure 2It is an optional flowchart of the industrial scenario recognition method provided by an embodiment of this application. Figure 2 The method in Figure 2 may include but is not limited to the following steps 101 to 104. When the industrial scenario recognition device executes the industrial scenario recognition method, the specific process is as follows. It should be noted first that this embodiment does not
[0071] Step 101: Obtain an industrial video stream in a target industrial scenario. Among them, there are multiple industrial objects in the target industrial scenario, and the industrial video stream includes multiple industrial images collected from the target industrial scenario from a past moment to the current moment.
[0072] The following provides a detailed description of step 101.
[0073] Among them, an industrial scenario refers to a specific industrial production environment or operation area. Exemplarily, an assembly line in an automobile manufacturing factory, a semiconductor chip manufacturing workshop, or a logistics warehouse can all be regarded as different industrial scenarios. In different industrial scenarios, relevant personnel can carry out corresponding industrial production activities, and industrial equipment will also operate in different production states. Therefore, industrial scenarios cover a series of processes and activities from raw material processing, production manufacturing to finished product inspection. For example, in the industrial scenario of a logistics warehouse, a packaging worker (relevant personnel) is packing the goods conveyed by a conveyor belt (industrial equipment). The target industrial scenario refers to a specific industrial scenario to be recognized.
[0074] Among them, an industrial video stream refers to a series of video frames continuously obtained from an industrial environment through a camera or other video acquisition devices. These video frames are arranged in chronological order to form a real-time or historical video data stream. The industrial video stream can monitor the industrial production process in the industrial scenario. In addition, an industrial image is a single frame image in the industrial video stream, which captures all visible information of the industrial scenario at a certain moment. The visible information includes the states of various industrial objects at the current moment and industrial production activities.
[0075] Furthermore, recognizing an industrial scenario means recognizing the personnel operations, equipment changes, etc. involved in this process. By recognizing the industrial scenario monitored by the industrial video stream, it can assist relevant operators in quickly judging whether the current industrial production activities meet the established requirements or whether there are equipment failure problems, thereby improving industrial production efficiency.
[0076] Further, the acquisition device for obtaining the industrial video stream can be fixedly arranged at a certain position to collect industrial images within a specified frame and form an industrial video stream within a partial range of the target industrial scene; alternatively, the acquisition device can also be rotatably arranged at a certain position to randomly collect industrial images by dynamically changing the direction and form an industrial video stream within the entire range of the target industrial scene. Of course, only an example is given here, and the acquisition device can be set according to the actual situation, and the embodiments of the present application do not limit this.
[0077] Among them, the industrial object refers to a specific entity or element existing in the target industrial scene. The industrial object can be static (such as machines, equipment) or dynamic (such as operating personnel). Usually, there are multiple industrial objects in the target industrial scene. For example, in the industrial scene of a logistics warehouse, there are at least one packing worker and at least one conveyor belt. The specific reference of the industrial object changes with the change of the target industrial scene. Only an example is given here, and it does not mean that the present application limits this.
[0078] Step 102: Input the industrial video stream into a pre-trained scene recognition model, and recognize multiple industrial images in chronological order.
[0079] The following is a detailed description of step 102.
[0080] In some embodiments, once the industrial video stream is input into the pre-trained scene recognition model, the scene recognition model will perform single-frame independent recognition and multi-frame correlation recognition processing on multiple industrial images in the order of time occurrence. Among them, single-frame independent recognition refers to separately recognizing and analyzing each frame of industrial image, that is, only analyzing the image content according to a single frame; while multi-frame correlation recognition takes into account the relationship between consecutive frames and uses the information contained in multiple industrial images in the time series for more complex image analysis.
[0081] Further, before recognizing multiple industrial images, data preprocessing can be performed first. The data preprocessing includes but is not limited to image enhancement, noise removal, etc. Among them, image enhancement is used to perform operations such as size adjustment, cropping, and normalization on each frame of industrial image to improve the effect of subsequent processing; noise removal is used to remove noise in the image to ensure the accuracy of the features extracted subsequently.
[0082] Step 103: If there is a dynamic change in the industrial object in the industrial image at the current moment, determine the industrial image at the current moment as the target industrial image, and determine the comparison industrial image of the target industrial image at the adjacent moment.
[0083] The following is a detailed description of step 103.
[0084] Among them, the dynamic change of the industrial object means that there are obvious changes in the state of the same industrial object at the current moment and at the historical moment. Exemplarily, as Figure 3 shown, Figure 3 is an optional industrial image schematic diagram of the industrial scenario recognition method provided by the embodiment of the present application, Figure 3 showing the industrial image a at the historical moment and the industrial image b at the current moment. Among them, the industrial image a includes three blowers and a stationary flag, and the industrial image b includes three blowers, an operator, and a flying flag; it can be determined that there are obvious changes in the industrial image b compared with the industrial image a. Therefore, the industrial image b can be used as the target industrial image, and the industrial image a can be used as the comparison industrial image.
[0085] Among them, the adjacent moment can refer to the previous moment of the current moment, the next moment of the current moment, multiple consecutive moments before the current moment, or multiple consecutive moments after the current moment, and the industrial image at the adjacent moment is determined as the comparison industrial image. In this way, at least one comparison industrial image is used to assist in describing the target industrial image in the subsequent process, thereby improving the accuracy of industrial scenario recognition.
[0086] In some embodiments, if there is a dynamic change in the industrial object in the industrial image at the current moment, determining the industrial image at the current moment as the target industrial image includes the following steps:
[0087] (103.a.1) Select multiple first industrial images from the historical moments of the industrial video stream, where the image features of the multiple first industrial images are all the same at the preset positions;
[0088] (103.a.2) Perform inter-frame fusion processing on the multiple first industrial image frames to obtain a first updated image, and determine the first foreground information from the first updated image;
[0089] (103.a.3) Select a second industrial image from the current moment of the industrial video stream, and determine the second foreground information from the second industrial image;
[0090] (103.a.4) If the first foreground information and the second foreground information do not match, determine that there is a dynamic change in the second industrial image, and determine the second industrial image at the current moment as the target industrial image.
[0091] The following describes steps (103.a.1) to (103.a.4) in detail.
[0092] In some embodiments, the industrial image at the current moment is determined to be the second industrial image. Taking the current moment as the boundary, the moments earlier than the current moment are historical moments, and the industrial images corresponding to multiple historical moments in the industrial video stream are selected as the first industrial images. Among them, the image features of the selected multiple first industrial images at the preset positions are all the same. In other words, the scene states reflected by the first industrial images should be the same without significant changes or movements.
[0093] Among them, the preset positions are usually some relatively key positions in the industrial image, and they can be set according to the target industrial scene. For example, assuming that the target industrial scene is an assembly line in a certain automobile manufacturing factory, the preset positions can be a certain fixed area of the conveyor belt or the central workbench, etc.
[0094] Furthermore, the image features of the first industrial images at the preset positions can refer to pixel value features, color features, and texture features. For example, when the image feature is a color feature, the color features of the multiple first industrial images at the preset positions should be the same.
[0095] Exemplarily, for the target industrial scene of the assembly line in an automobile manufacturing factory, the industrial image at the current moment is selected as the target industrial image T; and several industrial images are selected from the past few minutes of the industrial video stream as the first industrial images t1, t2, t3, t4, and t5; when the preset position indicates a certain fixed area, the image features (such as pixel value features, color features, etc.) of the first industrial images t1, t2, t3, t4, and t5 in this fixed area should be basically the same to indicate that the target industrial scene has not changed during this period.
[0096] Furthermore, after performing inter-frame fusion on the multiple first industrial images, by removing transient interference factors (such as light flickering, minor jitters, etc.) and enhancing the image quality, multiple consecutive video frames are merged into a higher-quality or more stable first updated image. In this way, when there are slight light or device movement interferences in the multiple first industrial images, the industrial scene recognition device effectively suppresses these transient noises through inter-frame fusion to ensure that the scene recognition model first accurately knows the image representing static immobility in the target industrial scene, so that the scene recognition model can perform target industrial scene recognition based on the first updated image as a benchmark.
[0097] Among them, the inter-frame fusion can have the following specific implementation methods:
[0098] ① Average method: Perform pixel-level averaging on a series of consecutive frames to reduce noise and smooth the image;
[0099] ② Median filtering method: Use the median value of all frames at each pixel position for pixel-level averaging to reduce the influence of extreme values and better retain edge information;
[0100] ③ Weighted average method: Different frames may be assigned different weights for pixel-level averaging processing;
[0101] ④ Optical flow method: By analyzing the change of luminance patterns between adjacent frames to infer the movement direction and speed of an object, to guide the inter-frame fusion process.
[0102] It should be noted that the specific inter-frame fusion method adopted by the industrial scenario recognition device can be set according to the actual situation. The above is only an example for illustration and does not represent a limitation in this embodiment of the present application.
[0103] Among them, by extracting information significantly different from the background part from the first updated image, the first foreground information is obtained. Common methods for extracting the first foreground information include background subtraction, threshold segmentation, edge detection, etc.; then, the second foreground information is extracted from the second industrial image by using a method similar to the method for extracting the first foreground information; if it is found during the comparison process that the second foreground information does not match the first foreground information, then the second industrial image where the second foreground information is located is determined as the target industrial image.
[0104] Exemplarily, the first foreground information a1 and the second foreground information a2 are respectively extracted from A1 in the first industrial image and A2 in the second industrial image. The first foreground information a1 is characterized as a conveyor belt, and the second foreground information a2 is characterized as a conveyor belt and the stacked materials on the conveyor belt. That is, the first foreground information a1 and the second foreground information a2 do not match, then the second industrial image A2 can be determined as the target industrial image.
[0105] It can be understood that the background of the target industrial scene includes stable and unchanging parts, and the scene recognition model proposed in this embodiment of the present application compares the foreground information of at least two industrial images to quickly locate the target industrial image from the industrial video stream, thereby improving the real-time performance of recognizing the target industrial scene shown in the target video stream.
[0106] Step 104, extract the first image feature of the target industrial image, and extract the second image feature of the comparison industrial image, and perform scene description on the target industrial image based on the difference between the first image feature and the second image feature, and output the scene description information of the target industrial scene.
[0107] The following gives a detailed description of step 104.
[0108] Among them, the scene description information is a text description generated after analyzing the feature differences between the target industrial image and the comparison industrial image. Exemplarily, as Figure 4 shown, Figure 4It is a schematic diagram of an optional target industrial image and its scene description information provided by an embodiment of the present application. An industrial video stream containing the Figure 4 shown target industrial image is input into the scene recognition model trained by the embodiment of the present application. After processing the industrial video stream, the scene recognition model outputs Figure 4 the shown scene description information; it should be noted that Figure 4 from (M1) to (M4) in are all complete descriptions of the target industrial image. The industrial scene recognition device provides more comprehensive and detailed information from multiple angles by outputting multiple descriptions, so as to help relevant operators or relevant intelligent systems better understand the situation in the current industrial scene in the follow-up.
[0109] It can be understood that traditional industrial scene recognition methods only perform image recognition on single-frame images. A significant difference between industrial scenes and other scenes is that they are always in frequent changes (such as lighting, production brightness, etc.). The industrial scene recognition device proposed by the embodiment of the present application extracts and analyzes features of both the target industrial image and the comparison industrial image to provide a stable basic reference point by using the comparison industrial image, and then compares the differences between the first image feature and the second image feature to describe the scene of the target industrial image, thereby enhancing its adaptability to complex and changeable industrial scenes.
[0110] In some embodiments, the first image feature of the target industrial image is extracted, and the second image feature of the comparison industrial image is extracted. Based on the differences between the first image feature and the second image feature, the scene of the target industrial image is described, and the scene description information of the target industrial scene is output, including the following steps:
[0111] (104.a.1) Perform feature encoding processing on the target industrial image to obtain a first encoded feature, and perform feature encoding processing on the comparison industrial image to obtain a second encoded feature;
[0112] (104.a.2) Perform feature decoding processing on the first encoded feature to obtain the first image feature, and perform feature decoding processing on the second encoded feature to obtain the second image feature;
[0113] (104.a.3) Based on the second image feature, generate comparison description information of the comparison industrial image, and based on the first image feature and the comparison description information, determine the scene description information of the target industrial image.
[0114] The following will describe steps (104.a.1) to (104.a.3) in detail.
[0115] Among them, the scene recognition model proposed in the embodiments of this application is provided with a Vision Transformer (ViT) module and a Long Short-Term Memory (LSTM) module. ViT is a model based on the Transformer architecture. It divides the input image into multiple small patches, and then embeds these patches into a high-dimensional space to form a series of tokens, thereby realizing the purpose of converting a two-dimensional image into one-dimensional sequence data, laying a foundation for subsequent generation of scene description information; LSTM effectively captures long-term dependencies in time series by introducing a gating mechanism (input gate, forget gate, output gate) to control the information flow.
[0116] In some embodiments, after determining the target industrial image and the comparison industrial image, taking the target industrial image as an example, use ViT to perform the following processing on the target industrial image:
[0117] ① Image chunking and embedding:
[0118] Assume that the target industrial image X ∈ R H×W×C , where H, W, and C respectively represent the height, width, and number of channels of the image. The image is divided into N non-overlapping image chunks, and the size of each image chunk is P×P, then there is:
[0119]
[0120] For example, if the size of the target industrial image is 224×224 and the image chunk size is 16×16, then N = 196.
[0121] ② Image chunk flattening and linear mapping:
[0122] Flatten each image chunk into a vector and convert it into a feature vector of a fixed length through linear mapping (fully connected layer):
[0123] x p = Linear(Flatten(X p ))
[0124] where Linear represents linear processing and Flatten represents flattening processing; X p represents the p-th image chunk, and x p ∈ R D , and D is the embedding dimension.
[0125] ③ Adding positional encoding:
[0126] Add learnable absolute positional encoding E pos :
[0127] z 0 = [x cls ; x 1 + E pos,1 ; x 2 + E pos,2 ; …; x N + E pos,N
[0128] wherein, x cls is a classification marker ([CLS] token) for aggregating global information; E pos,i is the position information of the i-th image patch; x 1 to x N (i = 1, 2......N) are used to represent the original feature representation of the i-th image patch.
[0129] ④ Feature encoding processing:
[0130] Among them, ViT is usually stacked by L layers of Transformer encoders, and each layer of encoder is composed of a multi-head self-attention mechanism (Multi-Head Self-Attention, MHSA) and a feed-forward neural network (Feed-Forward Network, FFN).
[0131] Furthermore, based on the multi-head self-attention mechanism, a linear mapping process is performed on the position encoding z l-1 to map the input into queries (Q), keys (K), and values (V):
[0132] Q = z l-1 W Q , K = z l-1 W K , V = z l-1 W V
[0133] Then, attention scores are calculated based on the attention function (Attention):
[0134]
[0135] Furthermore, the scene recognition model can dynamically focus on different features based on the attention scores.
[0136] Then, the input data is processed using a non-linear activation function (such as GELU):
[0137] FFN(x) = GELU(xW 1 + b 1 )W 2 + b 2
[0138] ⑤ Residual connection and layer normalization
[0139] After the multi - head self - attention and the feed - forward network, a residual connection and a normalization layer are added to enhance the expressive power of the output features:
[0140]
[0141] Among them, MHSA represents the processing of the multi - head attention mechanism, and FFN represents the processing of the feed - forward neural network.
[0142] ⑥ Encoded feature extraction:
[0143] After passing through the L - layer encoder, the output vector of the [CLS] token is regarded as the global feature representation of the target industrial image or the comparison industrial image, and the corresponding first encoded feature and second encoded feature are used as the image feature input of the long - short - term neural network module.
[0144] It should be noted that the second encoded feature is obtained by processing the comparison industrial image through a processing method similar to that of the target industrial image, and the specific processing process will not be elaborated here.
[0145] In some embodiments, the process of obtaining the first encoded feature by performing feature encoding on the target industrial image includes the following steps:
[0146] (A.1) Based on a preset first weight value, perform feature encoding on the second foreground information of the target industrial image to obtain a foreground first encoded feature;
[0147] (A.2) Based on a preset second weight value, perform feature encoding on other image information of the target industrial image except the second foreground information to obtain a background first encoded feature, where the first weight value is greater than the second weight value;
[0148] (A.3) Based on the foreground first encoded feature and the background first encoded feature, obtain the first encoded feature.
[0149] The following describes steps (A.1) to (A.3) in detail.
[0150] In some embodiments, when describing the scene of the target industrial image, the corresponding second foreground information in the target industrial image often includes more important details. Relatively speaking, the contribution degree of other image information (such as the image information corresponding to the background part) of the target industrial image except the second foreground information to the finally generated scene description information is weak. Therefore, in the process of using ViT to perform feature encoding processing on the target industrial image, a higher processing weight (the first weight value) can also be assigned to the second foreground information, and a second weight value can be assigned to other image information. In this way, the scene recognition device can perform feature encoding processing on different regions of the target industrial image with different degrees of attention, improving the generation efficiency of the scene description information while ensuring the accuracy of the finally generated scene description information.
[0151] Further, ViT can perform feature encoding processing on the second foreground information based on the first weight value to obtain the first foreground encoding feature; and perform feature encoding processing on other image information based on the second weight value to obtain the first background encoding feature; then, fuse the first foreground encoding feature and the first background encoding feature to obtain the first encoding feature.
[0152] Further, the fused first encoding feature can be obtained by a simple splicing method, or the first foreground encoding feature and the first background encoding feature can be spliced again based on the first weight value and the second weight value to obtain the first encoding feature. Here is only an example, and in actual situations, the fusion method of the first foreground encoding feature and the first background encoding feature can be determined according to specific situations, and the embodiments of the present application do not limit this.
[0153] It should be noted that the specific value of the first weight value can be set according to the actual situation, but usually the value of the first weight value is greater than the second weight value, so that when the scene recognition model performs encoding processing on the target industrial image, it pays more attention to the main part (the second foreground information) of the target industrial image.
[0154] In some embodiments, performing feature decoding processing on the first encoding feature to obtain the first decoding feature includes the following steps:
[0155] (B.1) Based on the first long short-term neural network provided with a tangent activation function, perform the first feature decoding processing on the first encoding feature to obtain the initial decoding feature;
[0156] (B.2) Based on the second long short-term neural network provided with a tangent activation function, perform the second feature decoding processing on the initial encoding feature to obtain the first decoding feature.
[0157] The following will describe steps (B.1) to (B.2) in detail.
[0158] In some embodiments, after completing the feature encoding process for the target industrial image and the comparison industrial image, ViT inputs the output result into the long short-term memory network module of the scene recognition model. Among them, the long short-term neural network module (MT-LSTM) includes a first long short-term neural network (the first layer of MT-LSTM) and a second long short-term neural network (the second layer of MT-LSTM). The encoded features are gradually decoded through the two-layer long short-term neural network to obtain the first decoded feature corresponding to the first encoded feature of the target industrial image. In this way, it is possible to more finely capture and reconstruct the feature information, thereby enhancing the overall expression ability of the finally output scene description information.
[0159] It should be noted that MT-LSTM is modified from the M-tanh activation function, and the specific formula of the M-tanh activation function is as follows:
[0160]
[0161] Among them, R is the set of real numbers; · is the element multiplication operator; the value range of H(x) is (-1 / m, 1 / m), centered at zero; the value range of the derivative of H(x) is (0, 1].
[0162] Furthermore, the M-tanh function is centered at zero, and the coefficient m can greatly accelerate the convergence speed, making it more suitable for the training and application of the long short-term memory network in the scene recognition device. Moreover, the value range of its derivative (0, 1] can effectively alleviate the vanishing gradient, and the coefficient 1 / m ensures that the derivative is not greater than 1, solving the problem of gradient explosion.
[0163] Furthermore, after determining the first encoded feature and the second encoded feature, taking the first encoded feature as an example, an explanation is given on how to use the long short-term neural network module to perform feature decoding processing on the first decoded feature:
[0164] ① Update of the candidate cell state:
[0165]
[0166] where x t is the input vector at the current moment, h t-1 is the hidden state at the previous moment, W c and b c respectively represent the weight and bias parameters of the long short-term neural network module, and tanh is the tangent activation function.
[0167] ② Update of the new cell state:
[0168]
[0169] where f tis the output of the forget gate, i t is the output of the input gate, C t-1 is the cell state at the previous moment, is the candidate cell state at the current moment.
[0170] Based on the above cell state update, the hidden state update of MT-LSTM can be expressed as:
[0171]
[0172] where, o t is the output of the output gate; C t is the updated cell state at the current moment; tanh is the tangent activation function; m is the parameter in the M-tanh activation function, which is used to control the convergence speed of the output.
[0173] Specifically, the embodiment of the present application also adds multi-head attention (Multi-headAttention) to the first layer of MT-LSTM, so that MT-LSTM can focus on different aspects of the input features from multiple different subspaces and improve the attention performance of the device.
[0174] Among them, the working principle of Multi-headAttention is to use the image descriptor I, the hidden state h at the previous moment t-1 , and the visual feature V as inputs, and calculate Q = h t-1 , K = V, V = V respectively. Introduce multiple attention heads (usually set to 8 heads), and perform the following operations on each head:
[0175] head i = Attention(QW i Q , KW i K , VW i V )
[0176] where, W i Q , W i K , W i V is the parameter matrix of the i-th attention head; head i is used to represent the i-th attention head.
[0177] Furthermore, the output results of all heads are concatenated and mapped through the linear layer W O to generate the final context vector C t . The generated context vector C tPassed to the hidden layer of MT-LSTM for state update. Specifically, the context vector C calculated using Multi-head Attention t replaces the context vector generated by the traditional Attention mechanism to determine the initial decoding feature; then, together with the input x t at the current time step t-1 and the hidden state h t are input into the second layer of MT-LSTM. The second layer of MT-LSTM will update its hidden state h t and cell state C t according to the context vector C t and the input x
[0178] to determine the first decoding feature (the first image feature). t-1 Furthermore, the input of the first layer of MT-LSTM is the current image descriptor I, the hidden state h t-1 of the current layer at the previous moment, the hidden state H t-1 of the second layer at the previous moment, and the output Y t of the second layer at the previous moment; this layer predicts the context vector C t through an adaptive attention mechanism with a visual sentinel (which helps the scene recognition model determine when to pay more attention to language rules and when to pay more attention to visual images), and predicts the value of G t through a block shift gate (when G i = 1, it means the current noun block ends, and the block pointer r t is moved to the next noun block; when G i = 0, it means the current noun block does not end, and r t is maintained at the current noun block). Then, the second layer of MT-LSTM takes C t and the hidden state h t of the previous layer at the current moment as inputs and outputs the predicted word Y
[0179] of the current layer; the output ends until the scene recognition model has predicted all the scene description information corresponding to the current target industrial scene.
[0180] Among them, a noun block refers to a semantically coherent noun phrase in a sentence. Since the second image feature corresponding to the comparison industrial image usually represents the known static situation of the target industrial scene, in the process of generating the scene description information corresponding to the target industrial image, considering the second image feature corresponding to the comparison industrial image on the basis of considering the first image feature corresponding to the target industrial image can help the scene recognition model better understand the target industrial image, thereby improving the accuracy of the finally generated scene description information and the output efficiency of the scene recognition model.It consists of it and its modifying components. Each noun block usually corresponds to a specific object or concept, such as equipment, tools, or operating states in industrial scenarios, like the equipment name "high-temperature furnace" or the industrial equipment "welding rod".
[0181] In some embodiments, the same method is used to obtain the corresponding second decoded features (second image features) of the comparison industrial image.
[0182] Exemplarily, assume that the target industrial image represents the current thermal image of a machine, and the comparison industrial image represents the thermal image of the machine in a normal state. After the scene recognition model respectively determines the first image features and the second image features corresponding to the target industrial image and the comparison industrial image, it also determines the corresponding comparison description information of the comparison industrial image based on the second image features. For example, the comparison description information is "This picture shows the machine in a normal operating state, and no abnormal hot spots are found."; Then, the scene recognition model will combine the first image features and the comparison description information to quickly determine the differential image features between the target industrial image and the comparison industrial image, and further obtain the scene description information. For example, the scene description information in this example can be "The target image shows a significant overheat spot in the motor area." It can be understood that if only a single-frame industrial image is analyzed and processed, the scene recognition model needs to analyze all the feature points in the image to determine the corresponding scene description information. The embodiments of the present application exactly take into account the drawbacks of the traditional method in this regard, and by making the scene recognition model compare and process multiple frames of images in the order of the time flow occurrence, the accuracy and efficiency of identifying the target industrial scene are improved.
[0183] In some embodiments, the scene recognition model is trained according to the following steps, and the steps include:
[0184] (104.b.1) Obtain the sample industrial video stream in the sample industrial scene and the corresponding sample label information of the sample industrial video stream. Among them, there are multiple sample industrial objects in the sample industrial scene, and the sample industrial video stream includes multiple frames of sample industrial images collected from the sample industrial scene from a past moment to the current moment;
[0185] (104.b.2) Input the sample industrial video stream and the sample label information into the pre-constructed initial scene recognition model, and recognize multiple frames of sample industrial images in chronological order;
[0186] (104.b.3) If there are dynamic changes in the sample industrial objects in the sample industrial image at the current moment, determine the sample industrial image at the current moment as the sample industrial image, and determine the sample comparison industrial image of the sample industrial image at the adjacent moment;
[0187] (104.b.4) Extract the sample first image feature of the sample industrial image, and extract the sample second image feature of the sample comparison industrial image. Describe the scene of the sample industrial image based on the difference between the sample first image feature and the sample second image feature, and output the sample scene description information of the sample industrial scene;
[0188] (104.b.5) Calculate the sample loss value representing the probability distribution difference between the sample scene description information and the sample label information, and calculate the sample evaluation value representing the description quality difference between the sample scene description information and the sample label information;
[0189] (104.b.6) Adjust the model parameters of the initial scene recognition model according to the sample loss value and the sample evaluation value. When the training stop condition is reached, the trained scene recognition model is obtained.
[0190] In some embodiments, as Figure 5 shown, Figure 5 FIG. is an optional schematic diagram for training a scene recognition model of the industrial scene recognition method provided by the embodiment of the present application. Input the sample industrial video stream and the sample label information into the pre-constructed initial scene recognition model, and output the predicted sample scene description information corresponding to the sample industrial video stream, so as to train the initial scene recognition model based on the sample scene description information and the corresponding sample label information later. Among them, the processing steps of the sample industrial video stream are similar to those in steps 101 to 104, and will not be elaborated here.
[0191] Among them, the model parameters can be weight values, bias values, activation function types, learning rates, batch sizes, specific values of coefficient m, hidden layer dimensions, number of attention heads, size of segmented image blocks, etc. Of course, the specific model parameters to be adjusted can be set according to the actual situation, and the embodiments of the present application do not limit this.
[0192] Among them, the training stop condition can be that the number of training epochs reaches a preset maximum value; or, when the change of the loss function used to evaluate the training degree is less than a certain threshold, it is considered that the initial image classification model has converged and the training stops; or, during the training process, an independent validation set is used to evaluate the performance of the model. If the performance on the validation set (such as accuracy, F1 score, etc.) no longer improves, the training stops. Of course, the training stop condition can be set according to the actual situation, and the embodiments of the present application do not limit this.
[0193] In some embodiments, calculating the sample loss value representing the probability distribution difference between the sample scene description information and the sample label information includes the following steps:
[0194] (C.1) Perform mapping processing on the sample label information to obtain sample label features;
[0195] (C.2) Calculate the probability distribution difference between the first decoded feature of the sample and the sample label feature based on a preset total loss function to obtain the sample loss value.
[0196] The following provides a detailed description of steps (C.1) to (C.2).
[0197] In some embodiments, as Figure 5 shown, the scene recognition model performs a mapping process on the input sample label information to obtain the sample label feature, and based on the total loss function L shown in the following formula total , determines the sample loss value:
[0198] L total = L caption + L gate
[0199] where
[0200]
[0201] In the formula, represents the given sample label information; represents the given set of regions; represents the value of the block shift gate corresponding to each time step; represents the selection probability of the current word, which is determined based on the first decoded feature; represents the transition probability of the current noun block; L caption (θ) represents the objective function for the rationality of the generated sentence; L gate (θ) represents the objective function for the consistency of the transition of noun blocks in the sentence with the input region sequence; When representing the cross-entropy loss function in a computer, torch.log() is used by default without writing the base number, and the log in the formula represents the natural logarithm function with base e.
[0202] In some embodiments, performing a mapping process on the sample label information to obtain the sample label feature includes the following steps:
[0203] (D.1) Obtain the industrial vocabulary corresponding to the sample industrial video stream, where the industrial vocabulary includes multiple industry-specific words;
[0204] (D.2) Compare the sample label information with the industry-specific words, and filter out redundant words from the sample label information to obtain the updated sample label information;
[0205] (D.3) Perform a mapping process on the updated sample label information to obtain the sample label feature.
[0206] The following provides a detailed description of steps (D.1) to (D.3).
[0207] In some embodiments, during the process of mapping sample label information, the pre-acquired industrial vocabulary is also used to preprocess the sample label information to improve the label quality, thereby enhancing the training efficiency of the initial scene recognition model.
[0208] Among them, the industrial vocabulary is a proprietary vocabulary constructed based on the characteristics of the target industrial scene domain, which contains high-frequency industrial feature words in the domain (such as "worker", "slag iron"). Each industrial feature word is mapped to a unique coding index. By restricting the input vocabulary range, it ensures that the scene recognition model focuses on domain-related semantics and improves the accuracy of understanding specific domain vocabulary.
[0209] Among them, redundant words refer to those contents that are irrelevant or repetitive to the output of the sample scene description information, which may interfere with the learning process of the initial scene recognition model and reduce the prediction accuracy. Further, the initial scene recognition model performs word segmentation on the input sample label information to obtain multiple basic words or sub-word units; then, the vocabulary is standardized, such as unifying case, handling synonyms and abbreviations, etc.; afterwards, redundant words (such as "the" or "are", etc.) are removed from them to obtain the key sample label features, which is convenient for improving the training efficiency of the initial scene recognition model in the subsequent process.
[0210] Further, as Figure 5 shown, the initial scene recognition model also performs word embedding processing on the sample label features corresponding to the sample label information. Specifically, based on the industrial vocabulary, each discrete sample label feature is mapped to a continuous high-dimensional vector representation, so that the generated embedding vector can reflect the semantic characteristics of the corresponding word and provide high-quality input features for subsequent processing.
[0211] In some embodiments, the sample label information is determined through the following steps, and the steps include:
[0212] (E.1) Based on the pre-set description template, different types of sample industrial objects in the sample industrial video stream are described to obtain the description information of each sample industrial object under the corresponding type, where the description templates corresponding to different types of sample industrial objects are different;
[0213] (E.2) Combine the corresponding description information of each sample industrial object to obtain the sample label information.
[0214] The following provides a detailed description of steps (E.1) to (E.2).
[0215] In some embodiments, as shown in Table 1 below, Table 1 is an exemplary description template that includes a unified specification for describing different types of sample industrial objects:
[0216] Table 1
[0217]
[0218] Furthermore, if multiple sample industrial objects are involved in the sample target industrial image, after completing the description of each type of sample industrial object, all the description information is merged to obtain the sample label information.
[0219] In some embodiments, calculating a sample evaluation value representing the description quality difference between the sample scenario description information and the sample label information includes the following steps:
[0220] (F.1) Parse the sample scenario description information and the sample label information respectively to obtain a plurality of first words and a plurality of second words;
[0221] (F.2) Based on a preset first evaluation function, calculate the sample evaluation value representing the matching degree between each first word and the corresponding second word;
[0222] (F.3) Alternatively, based on a preset second evaluation function, calculate the sample evaluation value representing the coverage degree between all the first words and all the second words.
[0223] The following provides a detailed description of steps (F.1) to (F.3).
[0224] In some embodiments, the sample evaluation value is calculated through the first evaluation function BLEU:
[0225]
[0226] Among them, BLEU is a commonly used indicator for evaluating the quality of generated text, mainly measuring the matching degree between the generated sample scenario description information and the sample label information at the lexical level. By comparing with the second words in the sample label information, the accuracy rate of the first word n-gram in the sample scenario description information is determined. Among them, BP is a penalty factor, p n is the precision value of the n-gram, w n is the corresponding weight value of the n-gram.
[0227] Alternatively, the sample evaluation value is calculated through the second evaluation function ROUGE-N:
[0228]
[0229] Among them, ROUGE is used to evaluate the recall rate in text generation tasks, measuring the coverage degree of the overlapping of n-gram between the generated sample scenario description information and the sample label information. The commonly used ROUGE metrics include ROUGE-N and ROUGE-L. Among them, gram n represents n-gram; Ref represents the sample label information, S represents a feature in the sample label information, and gram n represents a sequence of words with a continuous length of n; Count(gram n ) represents the number of matching n-grams.
[0230] As Figure 6 shown, Figure 6 is an optional flowchart of the industrial scenario recognition device provided by the embodiment of the present application. The industrial scenario recognition device includes the following modules 201 to 204:
[0231] An acquisition module 201, configured to acquire an industrial video stream in a target industrial scenario. Among them, there are multiple industrial objects in the target industrial scenario, and the industrial video stream includes multiple industrial images collected from the target industrial scenario from a past moment to the current moment;
[0232] An input module 202, configured to input the industrial video stream into a pre-trained scenario recognition model to recognize multiple industrial images in chronological order;
[0233] An industrial image determination module 203, configured to, if there is a dynamic change in the industrial object in the industrial image at the current moment, determine the industrial image at the current moment as the target industrial image, and determine the comparison industrial image of the target industrial image at an adjacent moment;
[0234] A target module 204, configured to extract the first image feature of the target industrial image, and extract the second image feature of the comparison industrial image, and perform a scenario description on the target industrial image based on the difference between the first image feature and the second image feature, and output the scenario description information of the target industrial scenario.
[0235] The industrial scenario recognition method, device, electronic device, and storage medium proposed in this application. The method obtains an industrial video stream in a target industrial scenario, where the target industrial scenario contains multiple industrial objects, and the industrial video stream includes multiple industrial images collected from the target industrial scenario from a past moment to the current moment; inputs the industrial video stream into a pre-trained scenario recognition model to recognize the multiple industrial images in chronological order; if there are dynamic changes in the industrial objects in the industrial image at the current moment, determines the industrial image at the current moment as the target industrial image, and determines the comparison industrial image of the target industrial image at an adjacent moment; extracts the first image feature of the target industrial image, and extracts the second image feature of the comparison industrial image, and describes the target industrial image based on the difference between the first image feature and the second image feature, and outputs the scene description information of the target industrial scenario. Since the second image feature corresponding to the comparison industrial image usually represents the known static situation of the target industrial scenario, on the basis of considering the first image feature corresponding to the target industrial image, the second image feature corresponding to the comparison industrial image is synchronously considered to help the scenario recognition model better understand the target industrial image, thereby improving the accuracy and real-time performance of the recognition of the target industrial scenario.
[0236] In some embodiments, a scenario recognition device including a trained scenario recognition model can be set in the target industrial scenario to achieve accurate and efficient recognition of the industrial production activities carried out in the target industrial scenario. The recognized scene description information can be directly sent to relevant operators, or sent to relevant intelligent systems for further processing. In practical applications, there is no need to consume a large amount of manpower to perform continuous monitoring operations. At the same time, the efficient and accurate scene description information can improve the automation level and safety of the industrial operation site.
[0237] The specific implementation manner of the industrial scenario recognition device is basically the same as the specific embodiments of the above industrial scenario recognition method, and will not be elaborated here.
[0238] The embodiments of this application also provide an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above industrial scenario recognition method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0239] As Figure 7 shown, Figure 7 is a schematic hardware structure diagram of the electronic device provided by the embodiments of this application. The electronic device includes:
[0240] The processor 301 can be implemented in the form of a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0241] The memory 302 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 302 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 302 and are called by the processor 301 to execute the industrial scenario recognition method of the embodiments of the present application;
[0242] The input / output interface 303 is used to implement information input and output;
[0243] The communication interface 304 is used to implement communication interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0244] The bus 305 transmits information between the various components of the device (such as the processor 301, the memory 302, the input / output interface 303, and the communication interface 304);
[0245] Among them, the processor 301, the memory 302, the input / output interface 303, and the communication interface 304 achieve communication connections with each other inside the device through the bus 305.
[0246] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned industrial scenario recognition method is implemented.
[0247] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories that are remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0248] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0249] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.
[0250] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0251] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0252] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above figures are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0253] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" may mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c may mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0254] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.
[0255] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0256] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0257] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0258] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of this application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of the rights of the embodiments of this application.
Claims
1. An industrial scene recognition method, characterized in that: include: Acquire an industrial video stream in a target industrial scene, wherein the target industrial scene contains a plurality of industrial objects, and the industrial video stream includes a plurality of frames of industrial images collected from the target industrial scene from a past moment to a current moment; Inputting the industrial video stream into a pre-trained scene recognition model, and recognizing multiple frames of the industrial images in time sequence; If the industrial object in the industrial image at the current moment has a dynamic change, the industrial image at the current moment is determined as a target industrial image, and a comparison industrial image of the target industrial image at an adjacent moment is extracted; A first image feature of the target industrial image is extracted, and a second image feature of the comparison industrial image is extracted, a scene description of the target industrial image is performed based on the difference between the first image feature and the second image feature, and scene description information of the target industrial scene is output.
2. The industrial scene recognition method according to claim 1, characterized in that: The industrial image includes a first industrial image and a second industrial image; If the industrial object in the industrial image at the current moment has a dynamic change, determining the industrial image at the current moment as a target industrial image includes: Selecting a plurality of the first industrial images from historical moments of the industrial video stream, wherein the image features of the plurality of the first industrial images at preset positions are consistent; Performing inter-frame fusion processing on a plurality of the first industrial image frames to obtain a first updated image, and determining first foreground information from the first updated image; Selecting the second industrial image from the current moment of the industrial video stream, and determining second foreground information from the second industrial image; If the first foreground information and the second foreground information do not match, it is determined that the second industrial image has a dynamic change, and the second industrial image at the current moment is determined as the target industrial image.
3. The industrial scene recognition method according to claim 2, characterized in that: The extracting the first image feature of the target industrial image and the extracting the second image feature of the comparison industrial image, performing scene description on the target industrial image based on the difference between the first image feature and the second image feature, and outputting scene description information of the target industrial scene includes: Performing feature coding processing on the target industrial image to obtain a first coding feature, and performing feature coding processing on the comparison industrial image to obtain a second coding feature; Performing feature decoding processing on the first coding feature to obtain a first image feature, and performing feature decoding processing on the second coding feature to obtain a second image feature; Based on the second image feature, comparative description information of the comparative industrial image is generated, and based on the first image feature and the comparative description information, the scene description information of the target industrial image is determined.
4. The industrial scene recognition method according to claim 3, characterized in that: The performing feature encoding processing on the target industrial image to obtain a first encoding feature includes: Based on a preset first weight value, performing feature coding processing on the second foreground information of the target industrial image to obtain a foreground first coding feature; Based on a preset second weight value, feature encoding processing is performed on other image information of the target industrial image except the second foreground information to obtain a background first encoding feature, wherein the first weight value is greater than the second weight value; The first coding feature is obtained based on the foreground first coding feature and the background first coding feature.
5. The industrial scene recognition method according to claim 3, characterized in that: The performing feature decoding processing on the first encoding feature to obtain a first decoding feature includes: Based on a first long short-term neural network provided with a tangent activation function, performing a first feature decoding process on the first encoding feature to obtain an initial decoding feature; Based on a second long short-term neural network provided with a tangent activation function, a second feature decoding process is performed on the initial coding feature to obtain the first decoding feature.
6. The industrial scene recognition method according to claim 1, characterized in that: The scene recognition model is trained according to the following steps, which include: Acquire a sample industrial video stream in a sample industrial scene, and sample label information corresponding to the sample industrial video stream, wherein the sample industrial scene contains a plurality of sample industrial objects, and the sample industrial video stream includes a plurality of frames of sample industrial images collected from the sample industrial scene from a past moment to a current moment; Inputting the sample industrial video stream and the sample label information into a pre-built initial scene recognition model, and recognizing multiple frames of the sample industrial images along the time sequence; If the sample industrial object in the sample industrial image at the current moment has a dynamic change, the sample industrial image at the current moment is determined as the sample industrial image, and a sample comparison industrial image of the sample industrial image at an adjacent moment is determined; Determine a sample first image feature of the sample industrial image, determine a sample second image feature of the sample comparison industrial image, perform a scene description on the sample industrial image based on a difference between the sample first image feature and the sample second image feature, and output sample scene description information of the sample industrial scene; Calculating a sample loss value representing a probability distribution difference between the sample scene description information and the sample label information, and calculating a sample evaluation value representing a description quality difference between the sample scene description information and the sample label information; According to the sample loss value and the sample evaluation value, the model parameters of the initial scene recognition model are adjusted until the training stop condition is reached, thereby obtaining the trained scene recognition model.
7. The industrial scene recognition method according to claim 6, characterized in that: The calculating and obtaining a sample loss value representing a probability distribution difference between the sample scene description information and the sample label information includes: Mapping the sample label information to obtain sample label features; Based on a preset total loss function, the probability distribution difference between the first decoding feature of the sample and the sample label feature is calculated to obtain the sample loss value.
8. The industrial scene recognition method according to claim 7, characterized in that: The mapping process is performed on the sample label information to obtain the sample label feature, including: Obtaining an industrial vocabulary corresponding to the sample industrial video stream, wherein the industrial vocabulary includes a plurality of industry-specific words; Comparing the sample label information with the industry-specific words, filtering out redundant words from the sample label information, and obtaining updated sample label information; Mapping processing is performed on the updated sample label information to obtain the sample label feature.
9. The industrial scene recognition method according to claim 6, characterized in that: The sample label information is determined by the following steps, which include: Based on a preset description template, different types of sample industrial objects in the sample industrial video stream are described to obtain description information of each sample industrial object under a corresponding type, wherein different types of sample industrial objects correspond to different description templates; The description information corresponding to each of the sample industrial objects is merged to obtain the sample label information.
10. The industrial scene recognition method according to claim 1, characterized in that: The calculating and obtaining a sample evaluation value characterizing the difference in description quality between the sample scene description information and the sample label information includes: Parsing the sample scene description information and the sample label information respectively to obtain a plurality of first words and a plurality of second words; Calculating the sample evaluation value representing the degree of matching between each of the first words and the corresponding second words based on a preset first evaluation function; Alternatively, based on a preset second evaluation function, the sample evaluation value representing the degree of coverage between all the first vocabularies and all the second vocabularies is calculated.
11. An industrial scene recognition device, characterized in that: include: An acquisition module, used for acquiring an industrial video stream in a target industrial scene, wherein the target industrial scene contains a plurality of industrial objects, and the industrial video stream includes a plurality of frames of industrial images collected from the target industrial scene from a past moment to a current moment; An input module, used to input the industrial video stream into a pre-trained scene recognition model, and recognize multiple frames of the industrial images along a time sequence; An industrial image determination module, configured to determine the industrial image at the current moment as a target industrial image if the industrial object in the industrial image at the current moment has a dynamic change, and extract a comparison industrial image of the target industrial image at an adjacent moment; The target module is used to extract the first image feature of the target industrial image and the second image feature of the comparison industrial image, perform scene description on the target industrial image based on the difference between the first image feature and the second image feature, and output scene description information of the target industrial scene.
12. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the industrial scene recognition method according to any one of claims 1 to 10 when executing the computer program.
13. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the industrial scene recognition method according to any one of claims 1 to 10 is implemented.