Image processing method, device, apparatus and storage medium
By performing multi-size segmentation and stitching of image frames, the problem of low accuracy in behavior recognition in existing technologies is solved, and more efficient feature detection and behavior recognition are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-28
- Publication Date
- 2026-03-31
AI Technical Summary
Existing behavior recognition methods do not have high enough accuracy.
The acquired image frames are segmented using grids of various sizes to obtain a set of image blocks. At least one image block in the set is then used for object recognition to obtain a set of target image blocks. Based on the position information of each target image block in the image frame, the blocks are stitched together to obtain an image to be analyzed containing a preset object. Finally, behavior recognition is performed on the image to be analyzed.
By using multi-size segmentation and stitching processing, more precise and detailed feature detection was achieved, improving the accuracy of behavior recognition.
Smart Images

Figure CN115482458B_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of computer vision, and in particular to an image processing method, apparatus, device, and storage medium. Background Technology
[0002] With the development of computer vision technology, behavior recognition based on computer vision has gradually been widely applied. For example, behavior recognition of videos or images collected in daily life scenarios (such as traffic intersections, residential communities, and subway lines) and production operation scenarios (such as workshops, construction sites, and high-risk work sites) can achieve intelligent monitoring of the corresponding scenarios, thereby ensuring the safety of people's lives and property as well as the safety of production operations. However, the behavior recognition methods in related technologies still have the problem of insufficient recognition accuracy. Summary of the Invention
[0003] In view of the above, embodiments of this application provide an image processing method, apparatus, device, and storage medium.
[0004] The technical solution of this application embodiment is implemented as follows:
[0005] On one hand, embodiments of this application provide an image processing method, the method comprising:
[0006] The acquired image frames are segmented using grids of various sizes to obtain a set of image blocks;
[0007] Object recognition is performed on at least one image patch in the image patch set to obtain a target image patch set; each target image patch in the target image patch set contains at least a portion of a preset object;
[0008] Based on the position information of each target image block in the image frame, at least one target image block is stitched together to obtain an image to be analyzed containing the preset object;
[0009] Behavior recognition is performed on the image to be analyzed to obtain the recognition result.
[0010] On the other hand, embodiments of this application provide an image processing apparatus, the apparatus comprising:
[0011] The segmentation module is used to segment the acquired image frames using grids of various sizes to obtain a set of image blocks;
[0012] The first recognition module is used to perform object recognition on at least one image block in the image block set to obtain a target image block set; each target image block in the target image block set contains at least a portion of a preset object;
[0013] The stitching module is used to stitch at least one of the target image blocks based on the position information of each target image block in the image frame to obtain an image to be analyzed containing the preset object;
[0014] The second recognition module is used to perform behavior recognition on the image to be analyzed and obtain the recognition result.
[0015] In another aspect, embodiments of this application provide a computer device, including a memory and a processor, wherein the memory stores a computer program that can run on the processor, and the processor executes the program to implement some or all of the steps in the above-described method.
[0016] In another aspect, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements some or all of the steps in the above-described method.
[0017] In this embodiment, image frames are segmented using grids of multiple sizes to obtain an image block set. Object recognition is performed on at least one image block in the image block set to obtain a target image block set. Each target image block in the target image block set contains at least a portion of a preset object. Based on the position information of each target image block in the image frame, at least one target image block is stitched together to obtain an image to be analyzed containing the preset object. Behavior recognition is then performed on the image to be analyzed to obtain a recognition result. Thus, on the one hand, by using grids of multiple sizes to segment the image frames and then performing object recognition on the resulting image blocks, more precise and detailed feature detection can be achieved while maintaining processing performance, thereby meeting the recognition requirements for objects of various sizes and improving the accuracy of behavior recognition. On the other hand, by stitching together at least one target image block based on its position in the image frame, the resulting image to be analyzed can contain more of the preset object or even the entire preset object. Therefore, more features of the preset object can be utilized during the behavior recognition process on the image to be analyzed, further improving the accuracy of behavior recognition. Attached Figure Description
[0018] Figure 1 A schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application;
[0019] Figure 2 A schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application;
[0020] Figure 3 A schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application;
[0021] Figure 4 A schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application;
[0022] Figure 5 A schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application;
[0023] Figure 6 A schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application;
[0024] Figure 7 A schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application;
[0025] Figure 8 A schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application;
[0026] Figure 9A This application provides a schematic diagram of a chemical tanker filling operation area in a semiconductor manufacturing process.
[0027] Figure 9B A schematic diagram illustrating a scenario where a fully enclosed fence is installed outside a interception ditch, as provided in an embodiment of this application;
[0028] Figure 9C A schematic diagram illustrating a scenario for setting up a fully enclosed fence, as provided in an embodiment of this application;
[0029] Figure 9D A schematic diagram illustrating a scenario where personnel within a fenced area are wearing Class C chemical protective suits, provided as an embodiment of this application.
[0030] Figure 9E A schematic diagram illustrating a scenario where a worker wears a safety belt or safety helmet, provided as an embodiment of this application.
[0031] Figure 9F A schematic diagram illustrating a scenario where wheel stops are provided for the drive wheels of a tank truck, as provided in an embodiment of this application;
[0032] Figure 9G A schematic diagram illustrating a scenario where personnel monitoring is present during a filling operation, as provided in an embodiment of this application;
[0033] Figure 9H This is a schematic diagram of the composition architecture of an image processing system provided in an embodiment of this application;
[0034] Figure 9I A schematic diagram illustrating a scenario in which multiple high-definition network cameras are installed in a chemical filling area, as provided in an embodiment of this application;
[0035] Figure 9JThis is a schematic diagram illustrating the implementation process of recognizing a preset object in an image processing method provided in an embodiment of this application.
[0036] Figure 9K A schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application;
[0037] Figure 10 This is a schematic diagram of the composition structure of an image processing device provided in an embodiment of this application;
[0038] Figure 11 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this application. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0040] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0041] If the application documents contain similar descriptions such as "first / second", the following explanation shall be added: In the following description, the terms "first / second / third" are used only to distinguish similar objects and do not represent a specific order of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0042] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this application.
[0043] This application provides an image processing method that can be executed by a processor of a computer device. The computer device refers to a device with data processing capabilities, such as a server, laptop, tablet, desktop computer, smart TV, set-top box, or mobile device (e.g., mobile phone, portable video player, personal digital assistant, dedicated messaging device, portable gaming device). Figure 1 This is a schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application, such as... Figure 1As shown, the method includes the following steps S101 to S104:
[0044] Step S101: The acquired image frames are segmented using grids of various sizes to obtain a set of image blocks.
[0045] Here, the image frame is the image to be used for behavior recognition. For example, it can be an image frame from a video stream or image sequence collected for everyday life scenarios (such as traffic intersections, residential communities, subway lines, etc.) or production operation scenarios (such as workshops, construction sites, high-risk work sites, etc.). In implementation, those skilled in the art can collect appropriate image frames according to actual needs, and this is not limited.
[0046] The various grid sizes can be preset or dynamically determined based on the size and resolution of the image frame; there is no limitation here.
[0047] Image frame segmentation refers to dividing an image frame into multiple image blocks, with each block set including at least one segmented block. In practice, the segmentation process can be performed multiple times using grids of different sizes based on the acquired image frame; the implementation method is not limited here.
[0048] In some embodiments, a first-size grid can be used to perform a first segmentation process on the acquired image frame, resulting in multiple image blocks. Then, a second-size grid can be used to perform a second segmentation process on the image blocks obtained from the first segmentation, and so on, until all grids of various sizes have been used to segment the image, resulting in a set of image blocks. In practice, grids can be selected sequentially in descending order of size for segmentation, or multiple segmentation processes can be performed using grids of various sizes in a random order; this is not limited. In some embodiments, after each segmentation process, the multiple image blocks can be identified and filtered to obtain at least one image block that meets set conditions. In the next segmentation process, only the at least one image block that meets the set conditions is segmented.
[0049] In some embodiments, the acquired image frames can be segmented using grids of various sizes, and an image block set is obtained based on the multiple image blocks obtained after segmentation by each grid size. In some embodiments, the multiple image blocks obtained after segmentation by each grid size can be identified and filtered to obtain at least one image block that meets the set conditions, and the at least one image block that meets the set conditions is added to the image block set.
[0050] Step S102: Perform object recognition on at least one image block in the image block set to obtain a target image block set; each target image block in the target image block set contains at least a portion of a preset object.
[0051] Here, the preset object is a pre-defined object containing behavior recognition-related features. Those skilled in the art can set appropriate preset objects according to actual circumstances; there is no limitation here. For example, in identifying abnormal operational behaviors during the loading and unloading of hazardous chemicals, the preset object can include any suitable object where violations may occur at the work site, including but not limited to one or more of personnel, chemical protective suits, safety helmets, safety belts, transport vehicles, and isolation barriers. The target image block contains at least a portion of the preset object; that is, the target image block contains the whole or a part of the preset object. For example, if the preset object includes personnel, the target image block can be an image block containing the complete human body, or an image block containing parts of the human body (such as the head, shoulders, arms, or lower legs); if the preset object includes chemical protective suits, the target image block can be an image block containing the entire chemical protective suit, or an image block containing parts of the chemical protective suit (such as sleeves, collars, zippers, pockets, or trouser legs).
[0052] By performing object recognition on image patches within a set of image patches, it can be determined whether the identified image patch contains part or all of a preset object. This allows for the identification of at least one target image patch from the set of image patches, resulting in a target image patch set. In implementation, any suitable image recognition algorithm can be used for object recognition of the image patches; there are no limitations. For example, an object detection algorithm can be used to detect the preset object in the image patch to determine whether the image patch contains part or all of the preset object. Alternatively, a classification algorithm can be used to classify the image patch, and the classification result can be used to determine whether the image patch is a target image patch.
[0053] Step S103: Based on the position information of each target image block in the image frame, at least one target image block is stitched together to obtain an image to be analyzed containing the preset object.
[0054] Here, since the target image blocks are segmented from the acquired image frames, each target image block has corresponding positional information within the image frame. During implementation, the positional information for each image block can be determined during the segmentation process.
[0055] The position information of an image block within an image frame can be any suitable information that characterizes the position of the image block within the image frame. In implementation, those skilled in the art can use appropriate position information to describe the position of the image block within the image frame according to the actual situation; this is not limited. For example, the position information of the image block within the image frame can be the coordinates of the four corners of the image block within the image frame, or it can be the coordinates of the center point of the image block.
[0056] Based on the positional information of each target image patch within the image frame, at least one target image patch can be stitched together, thereby combining adjacent target image patches to obtain at least one image to be analyzed. In this way, multiple target image patches containing parts of the same preset object and located adjacently within the image frame can be stitched together, resulting in an image to be analyzed that can contain more of the preset object or even the entire preset object. For example, when the preset object is a person, multiple target image patches containing human body parts can be stitched together based on their positional information within the image frame to obtain an image to be analyzed that contains a more complete picture of the person.
[0057] Step S104: Perform behavior recognition on the image to be analyzed to obtain the recognition result.
[0058] Here, the algorithm for behavior recognition in the image to be analyzed and the recognition results can be determined based on the actual application scenario; there are no limitations. For example, in the scenario of monitoring abnormal work behavior, any suitable behavior recognition algorithm can be used to identify abnormal work behavior in the image to be analyzed. The recognition results may include, but are not limited to, whether abnormal work behavior occurs in the current work scenario and the type of abnormal work behavior. Similarly, in the scenario of monitoring dangerous traffic behavior at an intersection, any suitable behavior recognition algorithm can be used to monitor dangerous traffic behavior in the image to be analyzed. The recognition results may include, but are not limited to, whether dangerous traffic behavior occurs in the current traffic scenario, the type of dangerous traffic behavior, and the license plate number of the vehicle involved.
[0059] In some embodiments, a pre-trained behavior recognition model can be used to perform behavior recognition on the image to be analyzed, and the recognition result can be obtained.
[0060] In this embodiment, image frames are segmented using grids of multiple sizes to obtain an image block set. Object recognition is performed on at least one image block in the image block set to obtain a target image block set. Each target image block in the target image block set contains at least a portion of a preset object. Based on the position information of each target image block in the image frame, at least one target image block is stitched together to obtain an image to be analyzed containing the preset object. Behavior recognition is then performed on the image to be analyzed to obtain a recognition result. Thus, on the one hand, by using grids of multiple sizes to segment the image frames and then performing object recognition on the resulting image blocks, more precise and detailed feature detection can be achieved while maintaining processing performance, thereby meeting the recognition requirements for objects of various sizes and improving the accuracy of behavior recognition. On the other hand, by stitching together at least one target image block based on its position in the image frame, the resulting image to be analyzed can contain more of the preset object or even the entire preset object. Therefore, more features of the preset object can be utilized during the behavior recognition process on the image to be analyzed, further improving the accuracy of behavior recognition.
[0061] This application provides an image processing method that can be executed by a computer device's processor. For example... Figure 2 As shown, the method includes:
[0062] Step S201: Acquire the image frame of the area to be identified.
[0063] Here, the area to be identified can be an area where behavior recognition is required. In implementation, the area to be identified can be the work area of a pre-set operation. For example, if the pre-set operation is a chemical filling operation, the area to be identified can be the chemical filling area; if the pre-set operation is a chemical unloading operation, the area to be identified can be the chemical unloading area; if the pre-set operation is a workshop production operation, the area to be identified can be the production workshop.
[0064] The method of acquiring image frames of the area to be identified is not limited in the embodiments of this application. In some embodiments, the computer device can acquire images of the area to be identified through an image acquisition module to obtain image frames of the area to be identified. In some embodiments, at least one camera can be set in the area to be identified, and after the camera acquires an image frame of the area to be identified, it can transmit the image frame to the computer device. In some embodiments, the computer device can also obtain stored image frames of the area to be identified from the Internet or a database.
[0065] Step S202: Determine whether a preset job exists in the image frame.
[0066] Here, the preset operation refers to a pre-defined operational process for which behavior recognition is to be performed. Those skilled in the art can determine the preset operation based on the actual scenario, and there is no limitation here. For example, the preset operation may include chemical filling operations, chemical unloading operations, or workshop production operations, etc.
[0067] In implementation, the presence of a pre-set operation in an image frame can be determined by considering the actual scenario and the characteristics of the pre-set operation; there are no limitations on this method. For example, for chemical filling operations, the presence of a chemical filling operation in an image frame can be determined by identifying the presence of vehicle features. If vehicle features are present, the chemical filling operation is confirmed; otherwise, it is not. Similarly, for workshop production operations, the presence of workshop production operations can be determined by identifying the presence of production equipment operating features in an image frame. If production equipment operating features are present, the workshop production operation is confirmed; otherwise, it is not. For example, during the process of performing a preset task, a sign indicating that the task is in progress can be set in the area to be identified. The presence of a preset task in an image frame can be determined by identifying whether the corresponding sign features exist in the image frame. If the sign features are found in the image frame, the preset task is confirmed to exist in the image frame. If the sign features are not found in the image frame, the preset task is confirmed to not exist in the image frame.
[0068] Step S203: If it is determined that there is a preset task in the image frame, the acquired image frame is segmented using grids of various sizes to obtain an image block set.
[0069] Step S204: Perform object recognition on at least one image block in the image block set to obtain a target image block set; each target image block in the target image block set contains at least a portion of a preset object.
[0070] Step S205: Based on the position information of each target image block in the image frame, at least one target image block is stitched together to obtain an image to be analyzed containing the preset object.
[0071] Step S206: Perform behavior recognition on the image to be analyzed to obtain the recognition result.
[0072] Here, steps S204 to S206 correspond to the aforementioned steps S102 to S104, and in practice, the specific implementation of the aforementioned steps S102 to S104 can be referred to.
[0073] In some embodiments, step S202 may include at least one of the following steps S221 and S222:
[0074] Step S221: Identify the features of the working equipment in the image frame. If the preset working equipment features exist in the image frame, determine that a preset job exists in the image frame.
[0075] Here, "operational equipment" refers to the equipment required to perform a pre-set operation. For example, for chemical filling or unloading operations, operational equipment may include transport vehicles used to transport chemicals, such as tank trucks and oil tankers. For workshop production operations, operational equipment may include production equipment used to produce products, such as machine tools and cutting machines.
[0076] The features of the work equipment refer to the characteristics related to the work equipment. Preset work equipment refers to features that characterize the presence of a preset work in an image frame. By identifying the features of the preset work equipment, it can be determined whether a corresponding work activity exists in the area to be identified. For example, if the work equipment includes a vehicle and the preset work includes chemical filling, the work equipment features can include features related to the vehicle. Since the presence of a vehicle in the image frame indicates the presence of a chemical filling operation, the preset work equipment features can include the vehicle's structural features, size, color, and movement pattern. Similarly, if the work equipment includes production equipment and the preset work includes workshop production operations, the work equipment features can include features related to the production equipment. Since the production equipment in the image frame is in operation, indicating the presence of workshop production operations, the preset work equipment features can include the production equipment's operating indicator lights, product generation frequency, and other features that characterize the production equipment's operational status.
[0077] In implementation, any suitable recognition algorithm can be used to identify the features of the working equipment in the image frame and determine whether the preset features of the working equipment exist in the image frame; there are no restrictions here.
[0078] Step S222: Identify the marking features in the image frame. If the image frame contains preset marking features, determine that the image frame contains a preset task.
[0079] Here, "marking features" refers to the relevant features of signs, banners, electronic screens, etc., used for marking. Preset marking features can be any suitable feature capable of indicating the start of a preset task or the progress of a preset task; there are no limitations. For example, before starting a preset task, markings indicating that a preset task is about to begin or is in progress (such as signs, banners, etc.) can be set in the area to be identified. Preset marking features can be features that characterize the presence of that marking in an image frame.
[0080] In implementation, any suitable recognition algorithm can be used to identify the marking features in the image frame to determine whether the preset marking features exist in the image frame; there are no restrictions here.
[0081] In some embodiments, step S221 may include steps S221a to S221d:
[0082] Step S221a: Obtain the video stream and extract a sequence of image frames containing multiple image frames from the video stream.
[0083] Here, the video stream can be captured by the image acquisition module of a computer device in the area to be identified, or it can be captured by a camera set in the area to be identified and transmitted to the computer device, or it can be downloaded by the computer device from the Internet or read from a database; there is no limitation here. In implementation, those skilled in the art can use appropriate methods to acquire the video stream according to the actual situation.
[0084] The image frame sequence may include multiple image frames of the region to be identified. These multiple image frames may be consecutive image frames in the video stream or multiple image frames with a specific frame interval; there are no limitations on this.
[0085] Step S221b: Identify the features of the operating equipment in the image frame sequence.
[0086] Here, the features of the operating equipment in multiple image frames of the region to be identified in the image frame sequence can be identified to determine whether there are image frames in the image frame sequence containing preset operating equipment features.
[0087] Step S221c: If it is determined that there is an image frame containing a preset operating equipment feature in the image frame sequence, determine the duration of the presence of the preset operating equipment feature in the video stream.
[0088] Here, if it is determined that there is an image frame containing the features of the preset operating equipment in the image frame sequence, the duration of the preset operating equipment features in the video stream can be determined by identifying the operating equipment features in multiple image frames before and after the image frame in the video stream.
[0089] Step S221d: If the duration exceeds the time threshold, determine that there is a preset job in the image frame sequence.
[0090] Here, the time threshold can be user-defined, a system default, or dynamically determined based on the type of preset task; there is no limitation. For example, the estimated duration of a preset task can be determined based on its type, and a time threshold proportional to that estimated duration can be determined accordingly.
[0091] In this embodiment, when a preset task is determined to exist in an image frame of the acquired region to be identified, the acquired image frame is segmented using grids of various sizes to obtain an image block set. This allows identification only of the behavior during the preset task, reducing unnecessary workload and improving the utilization of computing resources. In some embodiments, the presence of a preset task in an image frame can be quickly and easily determined by identifying the presence of preset task device features and / or preset identification features. In some embodiments, an image frame sequence containing multiple image frames can be extracted from a video stream. If an image frame containing preset task device features is found in the image frame sequence, and the presence duration of the preset task device features in the video stream exceeds a time threshold, the presence of a preset task in the image frame is determined. This improves the accuracy and detection rate of preset task identification, thereby improving the accuracy and detection rate of behavior recognition and further enhancing the utilization of computing resources.
[0092] This application provides an image processing method that can be executed by a computer device's processor. For example... Figure 3 As shown, the method includes:
[0093] Step S301: According to the order of grid size from large to small, at least one (i-1)th target image block is segmented using the i-th grid to obtain multiple i-th image blocks. Object recognition is performed on each i-th image block to obtain at least one i-th target image block that contains at least a portion of the preset object. Here, i is a positive integer less than N, the 0th target image block is the acquired image frame, and each i-th target image block contains at least a portion of the preset object.
[0094] Here, N is an integer greater than 1. The acquired image frames can be segmented using N different grid sizes, from largest to smallest, to obtain a set of image blocks. The i-th grid is the i-th grid among the N different grid sizes, from largest to smallest.
[0095] Step S302: When i is determined to be equal to N-1, the Nth grid is used to segment at least one (N-1)th target image block to obtain multiple Nth image blocks.
[0096] Step S303: Determine the image block set based on the multiple Nth image blocks.
[0097] Here, multiple Nth image blocks can be directly added to the image block set, or a set number of Nth image blocks can be selected from multiple Nth image blocks to add to the image block set according to the actual situation. There is no limitation here.
[0098] Step S304: Perform object recognition on at least one image block in the image block set to obtain a target image block set; each target image block in the target image block set contains at least a portion of a preset object.
[0099] Step S305: Based on the position information of each target image block in the image frame, at least one target image block is stitched together to obtain an image to be analyzed containing the preset object.
[0100] Step S306: Perform behavior recognition on the image to be analyzed to obtain the recognition result.
[0101] Here, steps S304 to S306 correspond to the aforementioned steps S102 to S104, and in practice, the specific implementation of the aforementioned steps S102 to S104 can be referred to.
[0102] In some embodiments, N is 3, and the N different sizes of the grids include a first grid, a second grid, and a third grid, wherein the size of the second grid is smaller than the size of the first grid, and the size of the third grid is smaller than the size of the second grid.
[0103] In this embodiment, according to the order of grid size from largest to smallest, at least one (i-1)th target image block is segmented using the i-th grid to obtain multiple i-th image blocks. Object recognition is then performed on each i-th image block to obtain at least one i-th target image block containing at least a portion of a preset object. If i equals N-1, at least one (N-1)th target image block is segmented using the N-th grid to obtain multiple N-th image blocks. Based on these multiple N-th image blocks, an image block set is determined. Thus, in the process of segmenting using multiple grids of decreasing size, smaller grids can be used to further segment the target image blocks obtained from larger grids, allowing for more precise and detailed object recognition on the further segmented image blocks. This results in a more accurate set of target image blocks, further improving the accuracy of behavior recognition. Furthermore, since smaller grids only further segment the target image blocks segmented by larger grids, more detailed object recognition is performed only on the image blocks segmented by smaller grids, reducing the workload of object recognition and improving image processing performance.
[0104] This application provides an image processing method that can be executed by a computer device's processor. For example... Figure 4 As shown, the method includes:
[0105] Step S401: The acquired image frames are segmented using grids of various sizes to obtain a set of image blocks.
[0106] Here, step S401 corresponds to the aforementioned step S101, and the specific implementation of the aforementioned step S101 can be referred to during implementation.
[0107] Step S402: For each image block in the image block set, a convolutional neural network is used to identify a preset object in the image block to obtain the probability that the image block contains the preset object.
[0108] Here, the convolutional neural network can be pre-trained using any suitable method. In implementation, those skilled in the art can use appropriate network structures to implement the convolutional neural network according to the actual situation, and there are no limitations here.
[0109] Step S403: Based on the probability that each of the image blocks contains the preset object, determine at least one target image block containing the preset object; each target image block in the set of target image blocks contains at least a portion of the preset object.
[0110] Here, whether an image block contains a preset object can be determined based on the probability that the image block contains the preset object. In implementation, image blocks with a probability greater than a set probability threshold can be identified as target image blocks. Alternatively, each image block in the image block set can be sorted in descending order of the probability of containing the preset object, and a predetermined number of image blocks at the top of the sorted list can be identified as target image blocks. Those skilled in the art can use appropriate methods to determine at least one target image block based on the probability of each image block containing the preset object, and the embodiments of this application are not limited in this regard.
[0111] Step S404: Determine a set of target image blocks based on the at least one target image block.
[0112] Here, at least one target image block can be directly added to the target image block set, or a set number of target image blocks can be selected from at least one target image block to add to the target image block set according to the actual situation. There is no limitation here.
[0113] Step S405: Based on the position information of each target image block in the image frame, at least one target image block is stitched together to obtain an image to be analyzed containing the preset object.
[0114] Step S406: Perform behavior recognition on the image to be analyzed to obtain the recognition result.
[0115] Here, steps S405 to S406 correspond to the aforementioned steps S103 to S104, and in practice, the specific implementation of the aforementioned steps S103 to S104 can be referred to.
[0116] In some embodiments, the number of preset objects is multiple, and step S402 may include the following steps S421 to S423:
[0117] Step S421: Perform convolution processing on the image patch to obtain an intermediate feature dataset.
[0118] Here, any suitable convolutional layer can be used to convolve the image patch according to the actual situation to extract features and obtain an intermediate feature dataset. The intermediate feature data contained in the intermediate feature dataset can be determined according to the actual convolutional layer used, and is not limited here.
[0119] Step S422: After performing data sampling and convolution processing on the intermediate feature dataset, a fully connected feature dataset is obtained.
[0120] Here, any suitable fully connected layer can be used to perform data sampling and convolution processing on the features in the intermediate feature dataset; there are no restrictions.
[0121] Step S423: Use an activation function to classify the features in the fully connected feature dataset to obtain the probability that the image patch contains different preset objects.
[0122] Here, the activation function can be one or more of the following, including but not limited to the sigmoid function, hyperbolic tangent (Tanh) function, rectified linear unit (RRC) function, and softmax function. In implementation, those skilled in the art can use any suitable activation function to classify the features in the fully connected feature dataset according to the actual situation, and obtain the probability that the image patch contains different preset objects; there is no limitation here.
[0123] In this embodiment, for each image patch in the image patch set, a convolutional neural network is used to identify a preset object in the image patch, obtaining the probability that the image patch contains the preset object. Then, based on the probability that each image patch contains the preset object, at least one target image patch containing the preset object is determined. Finally, based on at least one target image patch, a target image patch set is determined. In this way, object recognition can be performed on at least one image patch in the image patch set simply and quickly to obtain the target image patch set.
[0124] This application provides an image processing method that can be executed by a computer device's processor. For example... Figure 5 As shown, the method includes:
[0125] Step S501: The acquired image frames are segmented using grids of various sizes to obtain a set of image blocks.
[0126] Step S502: Perform object recognition on at least one image block in the image block set to obtain a target image block set; each target image block in the target image block set contains at least a portion of a preset object.
[0127] Here, steps S501 to S502 correspond to the aforementioned steps S101 to S102, and in implementation, the specific implementation of the aforementioned steps S101 to S102 can be referred to.
[0128] Step S503: Determine the type of each target image block based on the type of the preset object contained in each target image block.
[0129] Here, object recognition of image blocks can determine the type of preset objects contained in the target image block. The type of the target image block can characterize the type of preset objects contained in the target image block. Multiple target image blocks containing different types of preset objects can have the same type or different types, and this is not limited here. In some embodiments, target image blocks containing preset objects of the same or related types can be identified as belonging to the same class.
[0130] In implementation, those skilled in the art can use appropriate methods to classify the preset objects and target image blocks according to the actual situation, and determine the type of each target image block based on the type of the preset objects contained in each target image block. For example, in an abnormal operation behavior monitoring scenario, the type of preset objects may include one or more of personnel, chemical protective suits, safety helmets, safety belts, transportation vehicles, and isolation belts. Correspondingly, target image blocks including personnel, chemical protective suits, safety helmets, safety belts, transportation vehicles, and isolation belts can be determined as different types. In addition, since chemical protective suits, safety helmets, and safety belts are usually worn by people and are associated with personnel, target image blocks including personnel, chemical protective suits, safety helmets, and safety belts can also be determined as the same type, while target image blocks including transportation vehicles and isolation belts can be determined as other different types.
[0131] Step S504: Based on the position information of each target image block in the image frame, at least one target image block belonging to the same type is stitched together to obtain at least one image to be analyzed.
[0132] Here, based on the position information of each target image patch in the image frame, at least one target image patch of the same type is stitched together. This allows multiple adjacent target image patches of the same or related types of preset objects to be stitched together, resulting in each image to be analyzed containing more complete preset objects of the same or related types. For example, if the types of preset objects can include people, vehicles, and barriers, at least one target image patch containing people can be stitched together to obtain an image to be analyzed containing more complete people; at least one target image patch containing vehicles can be stitched together to obtain an image to be analyzed containing more complete vehicles; and at least one target image patch containing barriers can be stitched together to obtain an image to be analyzed containing more complete barriers.
[0133] Step S505: Perform behavior recognition on the image to be analyzed to obtain the recognition result.
[0134] Here, step S505 corresponds to the aforementioned step S104, and the specific implementation of the aforementioned step S104 can be referred to during implementation.
[0135] In this embodiment, the type of each target image block is determined based on the type of the preset object contained in each target image block, and at least one target image block of the same type is stitched together based on the position information of each target image block in the image frame to obtain at least one image to be analyzed. This allows for the separate identification of different preset objects in the image frame, thereby enabling more dimensional behavior recognition and further improving the accuracy and detection rate of behavior recognition.
[0136] This application provides an image processing method that can be executed by a computer device's processor. For example... Figure 6 As shown, the method includes:
[0137] Step S601: The acquired image frames are segmented using grids of various sizes to obtain a set of image blocks.
[0138] Step S602: Perform object recognition on at least one image block in the image block set to obtain a target image block set; each target image block in the target image block set contains at least a portion of a preset object, and the preset object includes preset transportation vehicle features.
[0139] Step S603: Based on the position information of each target image block in the image frame, at least one target image block is stitched together to obtain an image to be analyzed containing the preset object.
[0140] Here, steps S601 to S603 correspond to the aforementioned steps S101 to S103, and can be implemented with reference to the specific implementation of the aforementioned steps S101 to S103.
[0141] Step S604: Identify the preset transportation features in the image to be analyzed to determine whether the preset transportation features exist in the image to be analyzed.
[0142] Here, the preset transportation vehicle can be any suitable transportation vehicle determined according to the actual application scenario, including but not limited to tank trucks, oil tankers, and cement mixers. The preset transportation vehicle characteristics can be any suitable features that can be used to identify the preset transportation vehicle, such as the structural features, size, color, and movement pattern of the transportation vehicle, etc., and are not limited here.
[0143] In practice, those skilled in the art can use any suitable recognition algorithm to identify the preset transportation features in the image to be analyzed, and determine whether the preset transportation features exist in the image to be analyzed. This is not limited here.
[0144] Step S605: If it is determined that the preset transportation feature exists in the image to be analyzed, behavior recognition is performed on the image to be analyzed to obtain the recognition result.
[0145] In some embodiments, the preset object further includes an isolation zone feature, the recognition result includes a first recognition result, the number of images to be analyzed is at least one, and the behavior recognition of the images to be analyzed to obtain the recognition result in step S605 above includes the following steps S611 to S612:
[0146] Step S611: Identify the isolation band features in each of the images to be analyzed to determine whether the isolation band features exist in each of the images to be analyzed.
[0147] Here, the safety barrier can be any suitable object that can surround a pre-set transport vehicle in the area to be identified, such as red and white traffic cones, fences, etc., and is not limited here. The features of the safety barrier can be any suitable features that can be used to identify the safety barrier, and can be determined according to the actual object used for the safety barrier, and is not limited here.
[0148] In practice, those skilled in the art can use any suitable recognition algorithm to identify the isolation zone features in the image to be analyzed, and determine whether the isolation zone features exist in the image to be analyzed, which is not limited here.
[0149] Step S612: If it is determined that the isolation zone feature does not exist in any of the images to be analyzed, a first identification result is generated; the first identification result indicates that there is abnormal operation behavior in the area to be identified where no isolation zone is set.
[0150] In some embodiments, the recognition result further includes a second recognition result. The behavior recognition of the image to be analyzed described in step S605 above to obtain the recognition result further includes the following steps S621 to S622:
[0151] Step S621: If it is determined that the isolation zone feature exists in at least one of the images to be analyzed, the isolation zone feature in the at least one image to be analyzed is identified, and it is determined whether the target image block containing the isolation zone feature in the area to be identified surrounds the target image block containing the preset transportation feature.
[0152] Here, since each image to be analyzed is stitched together from at least one target image patch, for an image to be analyzed containing the features of a quarantine zone, the target image patch containing the quarantine zone features in the analyzed region can be determined by identifying the quarantine zone features in the analyzed image. For an image to be analyzed containing the features of a preset transportation vehicle, the target image patch containing the preset transportation vehicle features in the analyzed region can be determined by identifying the preset transportation vehicle features in the analyzed image. Based on the position information of each target image patch containing the quarantine zone features in the image frame, and the position information of each target image patch containing the preset transportation vehicle features in the image frame, it can be determined whether the target image patch containing the quarantine zone features in the region to be identified surrounds the target image patch containing the preset transportation vehicle features.
[0153] It should be noted that the image frames can be captured by at least one camera positioned around the area to be identified, and can include a 360-degree view of the area to be identified. In some embodiments, a single camera can capture the image of the area to be identified from a specific shooting angle, ensuring that the vehicle does not obstruct the barrier in the captured image frame, even if a pre-set vehicle and a safety barrier exist in the area to be identified. This reduces the likelihood of the barrier being obscured by the vehicle, thus allowing for accurate determination of whether a target image block containing barrier features surrounds a target image block containing pre-set vehicle features within the area to be identified. In other embodiments, multiple cameras can capture the image of the area to be identified from different shooting angles, obtaining a panoramic image frame containing the area to be identified, which can then accurately determine whether a target image block containing barrier features surrounds a target image block containing pre-set vehicle features within the area to be identified.
[0154] Step S622: If it is determined that the target image block corresponding to the isolation zone feature in the area to be identified does not surround the target image block corresponding to the preset vehicle feature, a second identification result is generated; the second identification result indicates that there is an abnormal operation behavior of a vehicle whose isolation zone is not fully enclosed in the area to be identified.
[0155] In some embodiments, the preset object further includes personnel characteristics, and the recognition result further includes a third recognition result. The behavior recognition of the image to be analyzed described in step S605 above to obtain the recognition result further includes the following steps S631 to S632:
[0156] Step S631: If it is determined that the isolation zone feature exists in at least one of the images to be analyzed, the isolation zone feature and personnel feature in the at least one image to be analyzed are identified to determine whether there is any violation by personnel in the area to be identified.
[0157] Here, personnel characteristics can be any suitable characteristics that can be used to identify personnel violations, determined according to the actual situation, including but not limited to human posture, clothing, and location. There are no restrictions on these characteristics.
[0158] Personnel violations can be any suitable behavior that does not conform to the personnel conduct norms or requirements of the area to be identified. In practice, those skilled in the art can determine the personnel conduct norms or requirements of the area to be identified based on the actual situation, and thus determine appropriate personnel violations; this is not limited. For example, if the area to be identified is a hazardous chemical loading and unloading area, workers within the isolation zone are required to wear specific protective clothing; therefore, personnel violations could include workers within the isolation zone not wearing protective clothing. Similarly, if the area to be identified is a construction site, all personnel on the construction site are required to wear safety helmets; therefore, personnel violations could include personnel on the construction site not wearing safety helmets.
[0159] Step S632: If it is determined that there is a violation by personnel in the area to be identified, a third identification result is generated; the third identification result indicates that there is a violation by personnel in the area to be identified.
[0160] In some embodiments, the preset object further includes features of chemical protective clothing, safety helmet, and safety belt, and the personnel violation includes at least one of the following: personnel located within the area enclosed by the isolation zone are not wearing chemical protective clothing, personnel located within the area to be identified are not wearing safety helmets, personnel located on the preset transport vehicle are not wearing safety belts, and there are no work supervisors in the area to be identified.
[0161] In implementation, for violations involving personnel within the isolation zone not wearing protective suits, at least one feature of the isolation zone, personnel, and protective suit in the image to be analyzed can be identified. If, within the area enclosed by a target image block corresponding to the isolation zone feature in the area to be identified, the target image block containing the personnel feature does not contain the protective suit feature, a characterization is generated indicating that personnel within the isolation zone in the area to be identified are violating regulations by not wearing protective suits. Here, the protective suit can be determined based on the actual application scenario of the area to be identified, and is not limited to any particular type. For example, if the area to be identified is a hazardous materials loading and unloading area, the protective suit can be a Class C protective suit that meets the requirements for hazardous materials operations.
[0162] For violations involving personnel not wearing helmets within the area to be identified, helmet features and personnel features in at least one image to be analyzed can be identified. If it is determined that a target image patch containing head features within the area to be identified does not contain helmet features, a characterizing violation of helmet-wearing by a person within the area to be identified is generated. Here, helmet features can be any suitable feature that can be used to identify helmets, and this embodiment of the application is not limited in this regard.
[0163] For the violation of not wearing a seatbelt by a person on a preset vehicle, features of the seatbelt, the person, and the preset vehicle in at least one image to be analyzed can be identified. If it is determined that a target image block containing features of the top of the preset vehicle is adjacent to the area to be identified, and the target image block containing the person does not contain seatbelt features, then a characterization is generated representing the existence of a violation of not wearing a seatbelt by a person on a preset vehicle in the area to be identified. Here, the seatbelt feature can be any suitable feature that can be used to identify a seatbelt; this embodiment of the application is not limited in this regard.
[0164] For violations by personnel without supervisory personnel within the area to be identified, at least one personnel feature and supervisory uniform feature in the image to be analyzed can be identified. If it is determined that none of the target image blocks containing personnel features within the area to be identified contain supervisory uniform features, a characterization is generated representing the presence of personnel violations without supervisory personnel within the area to be identified. Here, the supervisory personnel are those who oversee the work activities in the area to be identified, and these personnel may be wearing specific supervisory uniforms. The supervisory uniform feature can be any suitable feature that can be used to identify supervisory uniforms, such as the color and style of the uniform.
[0165] In some embodiments, the preset object further includes drive wheel features and wheel chock features, and the recognition result includes a fourth recognition result. The behavior recognition of the image to be analyzed described in step S605 above to obtain the recognition result includes the following steps S641 to S642:
[0166] Step S641: Identify the drive wheel feature and wheel chock feature in each of the images to be analyzed, and determine whether the target image block containing the drive wheel feature is connected to the target image block containing the wheel chock feature.
[0167] Here, the drive wheel feature can be any suitable feature used to identify the drive wheel of a vehicle. Since the drive wheel is typically the front wheel of a vehicle, in some embodiments, the drive wheel feature can be a feature of the front wheel of the vehicle, such as the positional features of the drive wheel.
[0168] Wheel chock features can be any suitable feature used to identify wheel chocks, such as the shape and material of the wheel chock.
[0169] In implementation, any suitable recognition algorithm can be used to identify the drive wheel features and wheel chock features in each image to be analyzed, and to determine whether the target image block containing the drive wheel features is connected to the target image block containing the wheel chock features. This application embodiment does not limit this.
[0170] Step S642: If it is determined that the target image block containing the drive wheel feature is not connected to the target image block containing the wheel chock feature, a fourth recognition result is generated; the fourth recognition result indicates that there is an abnormal operating behavior in the area to be recognized where the drive wheel is not equipped with a wheel chock.
[0171] In this embodiment, preset vehicle features are identified in the image to be analyzed to determine whether such features exist. If the preset vehicle features are found, behavior recognition is performed on the image to obtain recognition results. This allows for the identification of only the behavior during the loading and unloading process of the preset vehicle, thereby improving the utilization of computing resources. In some embodiments, the recognition results may include a first recognition result indicating abnormal operation behavior in the area to be identified (e.g., no guardrails), a second recognition result indicating abnormal operation behavior in the area to be identified (e.g., no guardrails), a third recognition result indicating personnel violations in the area to be identified, and / or a fourth recognition result indicating abnormal operation behavior in the area to be identified (e.g., no wheel chocks). This allows for the identification of abnormal operation behavior in the area to be identified from multiple dimensions, thereby improving the accuracy of abnormal operation behavior identification.
[0172] This application provides an image processing method that can be executed by a computer device's processor. For example... Figure 7 As shown, the method includes:
[0173] Step S701: The acquired image frames are segmented using grids of various sizes to obtain a set of image blocks.
[0174] Step S702: Perform object recognition on at least one image block in the image block set to obtain a target image block set; each target image block in the target image block set contains at least a portion of a preset object, and the preset object includes preset transportation vehicle features.
[0175] Step S703: Based on the position information of each target image block in the image frame, at least one target image block is stitched together to obtain an image to be analyzed containing the preset object.
[0176] Here, steps S701 to S703 correspond to the aforementioned steps S101 to S103, and can be implemented with reference to the specific implementation of the aforementioned steps S101 to S103.
[0177] Step S704: Identify the preset object in the image to be analyzed to obtain the location information of the key points in the image to be analyzed that correspond to the preset object.
[0178] Here, the location information of the key points corresponding to the preset object in the image to be analyzed refers to the location information of the preset object identified in the image frame. In practice, since the image to be analyzed is composed of target image blocks containing the preset object, and each target image block in the image to be analyzed is segmented from the image frame, the location information of each image block in the image frame can be determined during the segmentation process. Therefore, by identifying the preset object in the image to be analyzed, the target image blocks containing the preset object in the image to be analyzed can be determined, and thus the location information of the preset object in the image to be analyzed can be determined.
[0179] Step S705: Using the trained behavior recognition model, based on the location information of the key points, perform behavior recognition on the image to be analyzed to obtain the recognition result.
[0180] Here, those skilled in the art can use any suitable behavior recognition model according to the actual situation, and perform behavior recognition on the image to be analyzed based on the location information of key points. This application embodiment does not limit this.
[0181] For training the behavior recognition model, any suitable model training method can be used, and the embodiments of this application are not limited in this regard.
[0182] In some embodiments, prior to step S705 above, the method further includes:
[0183] Step S711: Obtain an abnormal work behavior sample set. Each sample image in the abnormal work behavior sample set has a category label and a location label. The category label is used to characterize the category of the abnormal work behavior corresponding to the sample image, and the location label is used to characterize the location information of the preset object in the sample image.
[0184] Here, the sample image can be any suitable image or image patch that can represent abnormal operating behavior. The category of abnormal operating behavior corresponding to the sample image refers to the abnormal operating behavior present in the scene corresponding to the sample image. The same sample image can correspond to one or more categories of abnormal operating behavior.
[0185] Abnormal work behaviors may include, but are not limited to, one or more of the following: abnormal work behaviors such as not setting up isolation zones in the area to be identified, abnormal work behaviors such as not fully enclosing the transport vehicle with isolation zones, personnel in the area enclosed by isolation zones not wearing chemical protective clothing, personnel in the area to be identified not wearing safety helmets, personnel in the pre-designated transport vehicle not wearing safety belts, no work supervision personnel in the area to be identified, and abnormal work behaviors such as not setting wheel chocks on the drive wheels of the pre-designated transport vehicle.
[0186] In practice, the category label and location label of each sample image in the abnormal operation behavior sample set can be pre-annotated manually or automatically by the system; there is no limitation on this.
[0187] Step S712: Use the abnormal operation behavior sample set to train the behavior recognition model to obtain the trained behavior recognition model.
[0188] Here, those skilled in the art can use any appropriate method to train the behavior recognition model using the abnormal operation behavior sample set according to the actual situation, and obtain a trained behavior recognition model; there is no limitation here.
[0189] In this embodiment, a preset object in the image to be analyzed is identified to obtain the location information of key points corresponding to the preset object in the image to be analyzed. Then, using a trained behavior recognition model, behavior recognition is performed on the image to be analyzed based on the location information of the key points, and the recognition result is obtained. Thus, because the location information of the preset object in the image frame is considered when performing behavior recognition on the image to be analyzed, the accuracy of behavior recognition can be further improved.
[0190] This application provides an image processing method that can be executed by a computer device's processor. For example... Figure 8 As shown, the method includes:
[0191] Step S801: The acquired image frames are segmented using grids of various sizes to obtain a set of image blocks.
[0192] Step S802: Perform object recognition on at least one image block in the image block set to obtain a target image block set; each target image block in the target image block set contains at least a portion of a preset object.
[0193] Step S803: Based on the position information of each target image block in the image frame, at least one target image block is stitched together to obtain an image to be analyzed containing the preset object.
[0194] Step S804: Perform behavior recognition on the image to be analyzed to obtain the recognition result.
[0195] Here, steps S801 to S804 correspond to the aforementioned steps S101 to S104, and can be implemented with reference to the specific implementation of the aforementioned steps S101 to S104.
[0196] Step S805: Based on the identification results, determine whether there is any abnormal operation behavior in the area to be identified.
[0197] Here, the identification results can represent the presence of abnormal operational behavior in the area to be identified in any suitable way; there are no limitations. Therefore, based on the identification results, it can be determined whether abnormal operational behavior exists in the area to be identified.
[0198] Step S806: If it is determined that there is abnormal operation behavior in the area to be identified, generate and send alarm information.
[0199] Here, alarm information refers to information used to alert to abnormal work behavior in the area to be identified. It may include, but is not limited to, one or more of the following: voice alarm information, alarm indicator light information, alarm telephone, alarm email, instant messaging software information, etc.
[0200] In some embodiments, generating and sending alarm information in step S806 above includes: step S811, determining the type of the abnormal operation behavior based on the identification result; and step S812, generating and sending alarm information based on the type of the abnormal operation behavior. Here, different alarm information can be generated for different types of abnormal operation behaviors, or the same alarm information can be generated. Alarm information can be sent in different ways, or in the same way; there is no limitation here.
[0201] In this embodiment, the system determines whether abnormal work behavior exists in the area to be identified based on the identification results. If abnormal work behavior is found in the area to be identified, an alarm message is generated and sent. In this way, by sending the alarm message, abnormal work behavior in the area to be identified can be detected in a timely manner, and the detected abnormal work behavior can be corrected promptly.
[0202] The following example, using the scenario of analyzing the behavior of supply personnel and environmental safety during the chemical tanker filling operation in semiconductor manufacturing, further illustrates the image processing method provided in this application embodiment.
[0203] Various chemicals are frequently used in semiconductor manufacturing processes, and these chemicals are stored, transported, and tanked by specialized personnel. For highly hazardous chemicals, such as tetramethylammonium hydroxide (TMAH), alcohol, and liquefied petroleum gas (LPG), serious safety accidents can occur if operators fail to follow proper procedures during tanker filling. Current technologies that rely solely on manual supervision of the tanker filling process are prone to oversights, making effective oversight of chemical filling difficult.
[0204] In view of this, embodiments of this application provide an image processing system capable of analyzing the behavior of supply personnel and environmental safety during the chemical tanker filling operation in semiconductor manufacturing processes based on deep learning.
[0205] Figure 9A This application provides a schematic diagram of a chemical tanker filling operation area in a semiconductor manufacturing process, as illustrated in the embodiments of this application. Figure 9A As shown, during the chemical tanker filling operation, operator 11 needs to fill the tanker 12 with chemicals. During the operation, a fully enclosed fence 13 needs to be set up around the tanker 12 to control personnel entering the work area and isolate non-operational personnel outside the chemical filling work area. In addition, monitoring personnel 14 are required in the chemical filling area to monitor the entire filling operation. Both the operator 11 and the monitoring personnel 14 in the chemical filling area wear safety helmets, and the operator 11 on the tanker also needs to wear a safety belt.
[0206] The image processing system provided in this application embodiment can intelligently monitor the following six major violations during the semiconductor chemical filling process:
[0207] 1) The chemical filling area is not fully enclosed; see [link / reference] Figure 9B A fully enclosed fence 13 may be set up outside the interception ditch 15. Failure to set up a fully enclosed fence 13 is a violation.
[0208] 2) Personnel not wearing Level C protective suits enter the area enclosed by the fence; see also Figure 9C Personnel not wearing Class C protective suits are not allowed to cross fence 13 to enter the area enclosed by the fence. Any entry of personnel not wearing Class C protective suits into the area enclosed by the fence is a violation.
[0209] 3) Personnel within the fenced area were not wearing Level C protective suits; see also Figure 9D Personnel within the fenced area are required to wear Class C protective suits. Failure to wear Class C protective suits within the fenced area constitutes a violation.
[0210] 4) Workers in the chemical filling area were not wearing safety helmets, or workers on tank trucks were not wearing safety belts; see also Figure 9E Workers 11 in the chemical filling area are required to wear safety helmets 16, and workers 11 on the tank truck 12 are required to wear safety belts 17. It is a violation for workers in the chemical filling area not to wear safety helmets or for workers on the tank truck not to wear safety belts.
[0211] 5) The tanker truck's drive wheels were not equipped with wheel chocks; here, the front wheels of the tanker truck are generally considered the drive wheels. In practice, the placement of the drive wheels should be considered in conjunction with the road slope of the chemical filling area to prevent the stopped tanker truck from rolling forward or backward. See also Figure 9F During the filling operation, wheel stops 18 need to be installed on the drive wheels 121 of the tank truck. Failure to install wheel stops on the drive wheels of the tank truck is a violation.
[0212] 6) There was no personnel monitoring the filling operation. See also Figure 9G In the chemical filling area, 14 monitoring personnel are required to monitor the entire filling operation. It is a violation if no personnel are monitoring the filling operation.
[0213] In the image processing system provided in this application embodiment, once at least one of the above-mentioned violations occurs in the chemical filling area, the type of violation can be automatically identified and an alarm message can be sent.
[0214] This application provides an image processing system that can analyze the behavior of supply personnel and environmental safety during the chemical tanker filling operation in semiconductor manufacturing. Figure 9H This is a schematic diagram of the composition architecture of an image processing system provided in an embodiment of this application, such as... Figure 9HAs shown, the system includes a high-definition network camera 21, a switch 22, a network video recorder (NVR) 23, a system server 24, and user terminal devices 25. In implementation, the high-definition network camera 21 is used to capture images of the chemical filling area, obtaining a video stream of the chemical filling area. The video stream captured by the high-definition camera 21 can be transmitted to the NVR 23 and the system server 24 via the switch 22. The NVR 23 can store and back up the video stream. The system server 24 can use a deep learning model (such as a convolutional neural network) to identify the features of preset objects (such as personnel, Class C protective suits, safety helmets, safety belts, fences, etc.) in each image frame of the video stream and determine whether any violations have occurred in the chemical filling area. Once a violation is detected, specific alarm information is generated according to the type of violation and sent to the user terminal device 25 (such as an office computer, personal computer, mobile phone, etc.) via email / instant messaging software. The high-definition network camera 21, switch 22, and network video recorder (NVR) 23 are all part of the system. The Recorder (NVR) 23 and the system server 24 can communicate through a specific video transmission network 31. The system server and the user terminal device can communicate through an office network 32. The video transmission network 31 and the office network 32 can be wired or wireless networks, and are not limited here.
[0215] In the image processing system provided in this application embodiment, multiple high-definition network cameras 21 can be installed in the chemical filling area to monitor the chemical filling area 24 / 7 from multiple angles. Figure 9I This application provides a schematic diagram of a scenario where multiple high-definition network cameras are installed in a chemical filling area, as illustrated in the embodiments of this application. Figure 9I As shown, multiple high-definition network cameras 21 can be installed on both sides of the road 41 to collect images of the chemical filling area 42 from all directions.
[0216] This application provides an image processing method, which can be executed by a system server in the image processing system provided in this application. See also Figure 9JThis method can segment the image frames 51 obtained from the video stream using a grid to obtain multiple segmented image blocks 52. A deep convolutional neural network is then used to identify each image block, obtaining the class probability (i.e., the probability that the image block contains different preset objects). Based on the class probability of each image block, target image blocks containing preset objects are selected. Then, target image blocks containing the same category of preset objects are stitched together in situ according to their position in the image frame to obtain the image to be analyzed 53. The image to be analyzed 53 contains preset objects such as personnel, safety helmets, or safety belts. Multiple images to be analyzed can be obtained after stitching (e.g., different workers, fences, wheel chocks, etc., each corresponding to a sub-image to be analyzed). These images can be combined with images containing different types of preset objects to determine whether any violations exist in the image frame. For example, by stitching together target image blocks containing people, safety helmets, and Class C protective suits, an image to be analyzed containing the whole person can be created. By using a behavior recognition model, the behavior of the people in the image to be analyzed can be judged to determine whether the workers are wearing safety helmets or safety belts. Then, by recognizing the image to be analyzed containing fences, it can be determined whether the people in the area enclosed by the fence are wearing Class C protective suits.
[0217] Figure 9K This is a schematic diagram illustrating the implementation flow of an image processing method provided in an embodiment of this application. This method can be implemented by a system server. Figure 9K As shown, the method includes the following steps S901 to S905:
[0218] Step S901: Acquire image frames from the video stream and divide the image frames into multiple image blocks using a grid.
[0219] Here, image frames can be divided into multiple image blocks of different sizes using grids of various sizes. Based on image blocks of different sizes, more accurate and detailed target detection can be achieved while maintaining performance, thus meeting the recognition requirements for preset objects of various sizes.
[0220] In some embodiments, acquiring image frames from the video stream requires meeting at least one of the following conditions: 1) A tanker truck is detected in the video stream. Since there is a certain time interval between the appearance of the tanker truck and the start of the actual filling operation, a time threshold can be preset. If a tanker truck appears in the frame and its duration exceeds the time threshold, an image frame from the video stream is acquired, and violation identification begins. For example, if a tanker truck appears in the frame and its duration exceeds the time threshold, and the image frame contains personnel features but no fence features, a violation is determined, and an alarm message is sent; 2) When the operator sets up a sign at the start of the filling operation, if the sign is detected in the video stream, an image frame from the video stream is acquired, and violation identification begins.
[0221] Step S902: Use a convolutional neural network to perform object recognition on each image block to obtain the class probability corresponding to each image block.
[0222] Here, the class probability corresponding to the image patch includes the probability that the image patch contains different preset objects.
[0223] Step S903: Select target image blocks containing preset objects according to the class probability corresponding to each image block, and stitch them together in situ according to the position of the target image blocks in the image frame to obtain the image to be analyzed.
[0224] Here, multiple target image blocks containing the same category can be stitched together in situ based on the category of the preset objects contained in the target image block and the position of the target image block in the original image to obtain the image to be analyzed.
[0225] Step S904: Establish a behavior recognition model based on a convolutional neural network, and train the behavior recognition model using the violation training set to obtain a trained behavior recognition model.
[0226] Here, the violation training set includes a specific number of sample images, which can be any suitable image or image patch that can reflect abnormal operational behavior. During implementation, the location information of preset targets and the existing violations in each sample image can be labeled in advance.
[0227] Step S905: Obtain the location information of key points corresponding to preset objects in the image to be analyzed, and input the location information of key points into the trained personnel behavior recognition model to obtain the behavior recognition result.
[0228] In some embodiments, the class probability corresponding to an image patch is obtained by recognizing the image patch using a convolutional neural network, which can be achieved through the following steps S921 to S923:
[0229] Step S921: The image patch is convolved using depthwise separable convolution to extract features and form an intermediate feature dataset.
[0230] Step S922: After performing multiple data sampling and convolution processes on the intermediate feature dataset, a fully connected feature dataset is obtained.
[0231] Step S923 involves calculating features in the fully connected feature dataset and using the ReLU activation function to classify the features in the fully connected feature dataset, obtaining the class probabilities corresponding to the image patches. In practice, a pre-trained classification model can be used to classify the features in the fully connected feature dataset to obtain the class probabilities of the image patches.
[0232] In some embodiments, the behavior recognition model is trained using the violation training set to obtain a trained behavior recognition model, which can be achieved through the following steps S941 to S942:
[0233] Step S941: Obtain the violation training set and classify the sample images in the violation training set according to the violation category. In practice, the category of the violation corresponding to the sample image can be labeled manually, and the part of the sample image containing the preset object can be marked on the sample image by using a rectangle. The coordinate positions of the four vertices of the rectangle are extracted to confirm the position of the preset object in the sample image.
[0234] Step S942: Input each sample image, along with the category of the violation and the coordinates of the four corners of the preset object within the rectangle in the sample image, into the behavior recognition model for training to obtain the trained behavior recognition model.
[0235] In some embodiments, step S905 may include the following steps S951 to S952:
[0236] Step S951: Use an object detection algorithm to identify objects in the image to be analyzed, and obtain the location information of key points in the image to be analyzed that correspond to preset objects.
[0237] Step S952: Using the trained behavior recognition model, the position information of key points corresponding to each preset object in the image to be analyzed is converted into feature vectors, and the feature vectors are classified to obtain the behavior recognition results.
[0238] In some embodiments, image frames can be segmented using a first grid, a second grid, and a third grid with sequentially decreasing sizes. First, the first grid is used to segment the image frame, and object recognition is performed on each segmented image block to obtain at least one first target image block containing a preset object. Then, the second grid is used to segment each first target image block, and object recognition is performed on each segmented image block to obtain at least one second target image block containing the preset object. Finally, the third grid is used to segment each second target image block, and object recognition is performed on each segmented image block to obtain at least one third target image block containing the preset object. The third target image blocks are then stitched together in situ according to their positions within the image frame to obtain the image to be analyzed. This improves upon the problem of slow object recognition speed caused by directly using small-sized grids for image segmentation, and also avoids the problem of low object recognition accuracy caused by using only large-sized grids for image segmentation.
[0239] In some embodiments, during object recognition in the image to be analyzed, different preset objects can be identified in a specific recognition order. For example, first, it is identified whether there is a tank truck in the image to be analyzed. If a tank truck is found, it is then identified whether there is a fence in the image to be analyzed. If a fence is found, it is then identified whether there are workers in the image to be analyzed. Based on the positional relationship between the workers, the tank truck, and the fence, it is determined whether there are any violations by personnel or other violations.
[0240] It should be noted that, in practice, the violations in this application embodiment may correspond to the abnormal operating behaviors in the aforementioned embodiments, and the violation training set may correspond to the abnormal operating behavior sample set in the aforementioned embodiments.
[0241] In this embodiment, by analyzing the behavior of supply personnel and environmental safety during the chemical tanker filling process in semiconductor manufacturing, chemicals in the semiconductor process can be effectively regulated, thereby improving the safety level and intelligence level of the semiconductor factory. Furthermore, various significant violations can be identified, allowing for multi-dimensional monitoring of the chemical tanker filling process, further enhancing the safety level and intelligence level of the semiconductor factory.
[0242] Figure 10 This is a schematic diagram of the composition structure of an image processing device provided in an embodiment of this application, as shown below. Figure 10 As shown, the image processing device 1000 includes: a segmentation module 1010, a first recognition module 1020, a stitching module 1030, and a second recognition module 1040, wherein:
[0243] The segmentation module 1010 is used to segment the acquired image frames using grids of various sizes to obtain a set of image blocks;
[0244] The first recognition module 1020 is used to perform object recognition on at least one image block in the image block set to obtain a target image block set; each target image block in the target image block set contains at least a portion of a preset object;
[0245] The stitching module 1030 is used to stitch together at least one of the target image blocks based on the position information of each target image block in the image frame to obtain an image to be analyzed containing the preset object.
[0246] The second recognition module 1040 is used to perform behavior recognition on the image to be analyzed and obtain the recognition result.
[0247] In some embodiments, the apparatus further includes: a first acquisition module, configured to acquire image frames of the acquired region to be identified; a first determination module, configured to determine whether a preset task exists in the image frame; the segmentation module is further configured to: if a preset task is determined to exist in the image frame, segment the acquired image frame using grids of multiple sizes.
[0248] In some embodiments, the first determining module is further configured to include at least one of the following: identifying the features of the working equipment in the image frame, and determining that a preset job exists in the image frame if the preset working equipment features exist in the image frame; identifying the marking features in the image frame, and determining that a preset job exists in the image frame if the preset marking features exist in the image frame.
[0249] In some embodiments, the first determining module is further configured to: acquire a video stream, extract an image frame sequence containing multiple image frames from the video stream; identify the features of the operating equipment in the image frame sequence; if it is determined that there is an image frame containing a preset operating equipment feature in the image frame sequence, determine the duration of the presence of the preset operating equipment feature in the video stream; if the duration of presence exceeds a time threshold, determine that there is a preset job in the image frame sequence.
[0250] In some embodiments, the multiple sizes of the grids include N different sizes of grids, where N is an integer greater than 1. The segmentation module is further configured to: segment at least one (i-1)th target image block using the i-th grid in descending order of grid size to obtain multiple i-th image blocks, and perform object recognition on each i-th image block to obtain at least one i-th target image block containing at least a portion of the preset object; wherein i is a positive integer less than N, the 0th target image block is an acquired image frame, and each i-th target image block contains at least a portion of the preset object; when i is determined to be equal to N-1, segment at least one (N-1)th target image block using the N-th grid to obtain multiple N-th image blocks; and determine the image block set based on the multiple N-th image blocks.
[0251] In some embodiments, N is 3, and the N different sizes of the grids include a first grid, a second grid, and a third grid, wherein the size of the second grid is smaller than the size of the first grid, and the size of the third grid is smaller than the size of the second grid.
[0252] In some embodiments, the first recognition module is further configured to: for each image block in the image block set, use a convolutional neural network to recognize a preset object in the image block to obtain the probability that the image block contains the preset object; based on the probability that each image block contains the preset object, determine at least one target image block containing the preset object; and based on the at least one target image block, determine a target image block set.
[0253] In some embodiments, the number of preset objects is multiple, and the first recognition module is further configured to: perform convolution processing on the image block to obtain an intermediate feature dataset; perform data sampling and convolution processing on the intermediate feature dataset to obtain a fully connected feature dataset; and use an activation function to classify the features in the fully connected feature dataset to obtain the probability that the image block contains different preset objects.
[0254] In some embodiments, the stitching module is further configured to: determine the type of each target image block based on the type of a preset object contained in each target image block; and stitch together at least one target image block of the same type based on the position information of each target image block in the image frame to obtain at least one image to be analyzed.
[0255] In some embodiments, the preset object includes preset transportation features, and the second identification module is further configured to: identify the preset transportation features in the image to be analyzed, determine whether the preset transportation features exist in the image to be analyzed; and, if the preset transportation features are determined to exist in the image to be analyzed, perform behavior recognition on the image to be analyzed to obtain an identification result.
[0256] In some embodiments, the preset object further includes a barrier zone feature, the identification result includes a first identification result, the number of images to be analyzed is at least one, and the second identification module is further configured to: identify the barrier zone feature in each image to be analyzed, determine whether the barrier zone feature exists in each image to be analyzed; generate a first identification result when it is determined that the barrier zone feature does not exist in any image to be analyzed; the first identification result indicates that there is an abnormal operation behavior in the area to be identified where a barrier zone is not set.
[0257] In some embodiments, the identification result further includes a second identification result, wherein the second identification module is further configured to: identify the isolation zone feature in at least one of the images to be analyzed, and determine whether a target image block containing the isolation zone feature in the area to be identified surrounds a target image block containing the preset vehicle feature; generate a second identification result if it is determined that the target image block corresponding to the isolation zone feature in the area to be identified does not surround the target image block corresponding to the preset vehicle feature; the second identification result characterizes the abnormal operation behavior of a vehicle whose isolation zone is not fully enclosed in the area to be identified.
[0258] In some embodiments, the preset object further includes personnel features, and the identification result further includes a third identification result. The second identification module is further configured to: identify the isolation zone features and personnel features in the at least one image to be analyzed when it is determined that the isolation zone features exist in at least one of the images to be analyzed, and determine whether there is any violation by personnel in the area to be identified; generate a third identification result when it is determined that there is any violation by personnel in the area to be identified; the third identification result indicates that there is any violation by personnel in the area to be identified.
[0259] In some embodiments, the preset object further includes features of chemical protective clothing, safety helmet, and safety belt, and the personnel violation includes at least one of the following: personnel located within the area enclosed by the isolation zone are not wearing chemical protective clothing, personnel located within the area to be identified are not wearing safety helmets, personnel located on the preset transport vehicle are not wearing safety belts, and there are no work supervisors in the area to be identified.
[0260] In some embodiments, the preset object further includes drive wheel features and wheel chock features, and the recognition result includes a fourth recognition result. The second recognition module is further configured to: recognize the drive wheel features and wheel chock features in each of the images to be analyzed, and determine whether a target image block containing the drive wheel features is connected to a target image block containing the wheel chock features; if it is determined that a target image block containing the drive wheel features is not connected to a target image block containing the wheel chock features, generate a fourth recognition result; the fourth recognition result indicates that there is an abnormal operating behavior in the area to be identified where the drive wheel is not equipped with a wheel chock.
[0261] In some embodiments, the second recognition module is further configured to: recognize the preset object in the image to be analyzed, and obtain the location information of key points in the image to be analyzed corresponding to the preset object; and use a trained behavior recognition model to perform behavior recognition on the image to be analyzed based on the location information of the key points, and obtain a recognition result.
[0262] In some embodiments, the apparatus further includes: a second acquisition module, configured to acquire an abnormal work behavior sample set, wherein each sample image in the abnormal work behavior sample set has a category label and a location label, wherein the category label is used to characterize the category of the abnormal work behavior corresponding to the sample image, and the location label is used to characterize the location information of the preset object in the sample image; and a training module, configured to use the abnormal work behavior sample set to train the behavior recognition model to obtain the trained behavior recognition model.
[0263] In some embodiments, the apparatus further includes: a second determining module, configured to determine whether there is abnormal work behavior in the area to be identified based on the identification result; and a sending module, configured to generate and send alarm information when it is determined that there is abnormal work behavior in the area to be identified.
[0264] In some embodiments, the sending module is further configured to: determine the type of the abnormal operation behavior based on the identification result; and generate and send alarm information based on the type of the abnormal operation behavior.
[0265] The descriptions of the above device embodiments are similar to those of the above method embodiments, and have similar beneficial effects. For technical details not disclosed in the device embodiments of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0266] It should be noted that, in the embodiments of this application, if the above-described image processing method is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware and software combination.
[0267] Correspondingly, embodiments of this application provide a computer device, including a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement the steps in the above-described method.
[0268] Correspondingly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps in the above-described method.
[0269] Correspondingly, embodiments of this application provide a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied as a computer storage medium; in another optional embodiment, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0270] It should be noted that the descriptions of the above-described storage media, computer program products, and device embodiments are similar to the descriptions of the above-described method embodiments, and have similar beneficial effects. For technical details not disclosed in the embodiments of the storage media, computer program products, and devices of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0271] It should be noted that, Figure 11 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this application, such as... Figure 11 As shown, the hardware entity of the computer device 1100 includes: a processor 1101, a communication interface 1102, and a memory 1103, wherein:
[0272] Processor 1101 typically controls the overall operation of computer device 1100.
[0273] Communication interface 1102 enables computer devices to communicate with other terminals or servers via a network.
[0274] The memory 1103 is configured to store instructions and applications executable by the processor 1101, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 1101 and various modules in the computer device 1100. It can be implemented by flash memory or random access memory (RAM).
[0275] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above-described processes do not imply a sequential order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above-described embodiments are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0276] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0277] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0278] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0279] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0280] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0281] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.
[0282] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An image processing method, characterized by, The method comprises: segmenting the collected image frame by using multiple sizes of grids to obtain a set of image blocks; performing object recognition on at least one image block in the set of image blocks to obtain a set of target image blocks; each target image block in the set of target image blocks contains at least part of a preset object; based on the position information of each target image block in the image frame, splicing at least one target image block to obtain an image to be analyzed containing the preset object; performing behavior recognition on the image to be analyzed to obtain a recognition result; wherein the multiple sizes of grids include N different sizes of grids, N is an integer greater than 1, and segmenting the collected image frame by using multiple sizes of grids to obtain a set of image blocks comprises: in order of grid size from large to small, sequentially segmenting at least one (i-1)th target image block by using the ith grid to obtain a plurality of ith image blocks, and performing object recognition on each ith image block to obtain at least one ith target image block containing at least part of the preset object; wherein i is a positive integer less than N, the 0th target image block is the collected image frame, and each ith target image block contains at least part of the preset object; in a case where it is determined that i is equal to N-1, segmenting at least one (N-1)th target image block by using the Nth grid to obtain a plurality of Nth image blocks.
2. The method of claim 1, wherein, The method further comprises: obtaining an image frame of a to-be-recognized region collected; determining whether a preset task exists in the image frame; segmenting the collected image frame by using multiple sizes of grids comprises: in a case where it is determined that the preset task exists in the image frame, segmenting the collected image frame by using multiple sizes of grids.
3. The method of claim 2, wherein, The determination of whether the preset task exists in the image frame comprises at least one of: recognizing a task equipment feature in the image frame, and in a case where the preset task equipment feature exists in the image frame, determining that the preset task exists in the image frame; recognizing an indication feature in the image frame, and in a case where the preset indication feature exists in the image frame, determining that the preset task exists in the image frame.
4. The method of claim 3, wherein, The recognition of the task equipment feature in the image frame, and the determination that the preset task exists in the image frame in a case where the preset task equipment feature exists in the image frame, comprises: obtaining a video stream, and cutting an image frame sequence containing a plurality of image frames from the video stream; recognizing a task equipment feature in the image frame sequence; in a case where it is determined that the image frame sequence contains an image frame containing the preset task equipment feature, determining the existing duration of the preset task equipment feature in the video stream; in a case where the existing duration exceeds a time threshold, determining that the preset task exists in the image frame sequence.
5. The method of claim 1, wherein, N is 3, and the N different sizes of grids include a first grid, a second grid, and a third grid, wherein the size of the second grid is smaller than the size of the first grid, and the size of the third grid is smaller than the size of the second grid.
6. The method according to any one of claims 1 to 4, characterized in that, The object recognition on at least one image block in the image block set obtains a target image block set, and the object recognition on the image block set comprises: For each image block in the image block set, a preset object in the image block is recognized by using a convolutional neural network to obtain a probability that the image block contains the preset object; At least one target image block containing the preset object is determined based on the probability that each image block contains the preset object; A target image block set is determined based on the at least one target image block.
7. The method of claim 6, wherein, The number of the preset objects is multiple, and the probability that the image block contains the preset object is obtained by recognizing the preset object in the image block by using the convolutional neural network, which comprises: The image block is subjected to convolution processing to obtain an intermediate feature data set; After data sampling and convolution processing on the intermediate feature data set, a fully connected feature data set is obtained; The features in the fully connected feature data set are classified by using an activation function to obtain the probability that the image block contains different preset objects.
8. The method according to any one of claims 1 to 4, characterized in that, The target image block set is determined based on the position information of each target image block in the image frame, and the at least one target image block is spliced to obtain an analysis image containing the preset object, which comprises: The type of each target image block is determined based on the type of the preset object contained in each target image block; At least one target image block belonging to the same type is spliced based on the position information of each target image block in the image frame to obtain at least one analysis image.
9. The method according to any one of claims 2 to 4, characterized in that, The preset object includes a preset transportation tool feature, and the behavior recognition on the analysis image obtains a recognition result, which comprises: The preset transportation tool feature in the analysis image is recognized to determine whether the preset transportation tool feature exists in the analysis image; In a case where it is determined that the preset transportation tool feature exists in the analysis image, the analysis image is subjected to behavior recognition to obtain a recognition result.
10. The method of claim 9, wherein, The preset object further includes a barrier feature, the recognition result includes a first recognition result, the number of the analysis images is at least one, and the behavior recognition on the analysis image obtains a recognition result, which comprises: The barrier feature in each analysis image is recognized to determine whether the barrier feature exists in each analysis image; In a case where it is determined that the barrier feature does not exist in each analysis image, a first recognition result is generated; the first recognition result represents that an abnormal operation behavior without setting a barrier exists in the to-be-recognized region.
11. The method of claim 10, wherein, The recognition result further includes a second recognition result, and the behavior recognition on the analysis image obtains a recognition result, which further comprises: In a case where it is determined that the barrier feature exists in at least one analysis image, the barrier feature in the at least one analysis image is recognized to determine whether a target image block containing the barrier feature in the to-be-recognized region surrounds a target image block containing the preset transportation tool feature; generate a second identification result in a case where it is determined that the isolation belt feature in the to-be-identified region does not correspond to a target image block surrounding the target image block corresponding to the preset transportation tool feature; the second identification result represents that there is an abnormal operation behavior of an isolation belt not fully enclosing a transportation tool in the to-be-identified region.
12. The method of claim 10, wherein, The preset object further includes a personnel feature, and the identification result further includes a third identification result, and the behavior identification on the to-be-analyzed image to obtain the identification result further includes: In a case where it is determined that the isolation belt feature exists in at least one of the to-be-analyzed images, identifying the isolation belt feature and the personnel feature in the at least one to-be-analyzed image to determine whether there is a personnel violation behavior in the to-be-identified region; In a case where it is determined that there is a personnel violation behavior in the to-be-identified region, generating a third identification result; the third identification result represents that there is a personnel violation behavior in the to-be-identified region.
13. The method of claim 12, wherein, The preset object further includes a chemical protection suit feature, a safety helmet feature, and a safety belt feature, and the personnel violation behavior includes at least one of the following: A personnel located in the region surrounded by the isolation belt does not wear a chemical protection suit, A personnel located in the to-be-identified region does not wear a safety helmet, A personnel located on the preset transportation tool does not wear a safety belt, There is no operation supervision personnel in the to-be-identified region.
14. The method of claim 9, wherein, The preset object further includes a driving wheel feature and a wheel stop feature, and the identification result includes a fourth identification result, and the behavior identification on the to-be-analyzed image to obtain the identification result further includes: identifying the driving wheel feature and the wheel stop feature in each of the to-be-analyzed images to determine whether a target image block containing the driving wheel feature is connected to a target image block containing the wheel stop feature; In a case where it is determined that the target image block containing the driving wheel feature is not connected to the target image block containing the wheel stop feature, generating a fourth identification result; the fourth identification result represents that there is an abnormal operation behavior of the driving wheel without a wheel stop in the to-be-identified region.
15. The method according to any one of claims 1 to 4, characterized in that, The behavior identification on the to-be-analyzed image to obtain the identification result includes: identifying the preset object in the to-be-analyzed image to obtain position information of a key point corresponding to the preset object in the to-be-analyzed image; using a trained behavior identification model to perform behavior identification on the to-be-analyzed image based on the position information of the key point to obtain the identification result.
16. The method of claim 15, wherein, Before the behavior identification on the to-be-analyzed image based on the position information of the key point to obtain the identification result using the trained behavior identification model, the method further includes: obtaining an abnormal operation behavior sample set, each sample image in the abnormal operation behavior sample set having a class label and a position label, the class label being used to represent a class of an abnormal operation behavior corresponding to the sample image, and the position label being used to represent position information of the preset object in the sample image; training the behavior identification model using the abnormal operation behavior sample set to obtain the trained behavior identification model.
17. The method of any one of claims 2 to 4, wherein, The method further includes: determining whether there is an abnormal operation behavior in the to-be-identified region based on the identification result; In a case where it is determined that the abnormal operation behavior exists in the to-be-identified region, alarm information is generated and sent.
18. The method of claim 17, wherein, The generating and sending of the alarm information comprises: Based on the identification result, the type of the abnormal operation behavior is determined; Based on the type of the abnormal operation behavior, the alarm information is generated and sent.
19. An image processing apparatus employing the image processing method according to any one of claims 1 to 18, characterized by Comprise: The segmentation module is used for segmenting the collected image frames by using multiple sizes of grids to obtain a set of image blocks; The first identification module is used for performing object identification on at least one image block in the set of image blocks to obtain a set of target image blocks; each target image block in the set of target image blocks at least contains part of a preset object; The splicing module is used for splicing at least one target image block based on the position information of each target image block in the image frame to obtain a to-be-analyzed image containing the preset object; The second identification module is used for performing behavior identification on the to-be-analyzed image to obtain an identification result.
20. A computer device, comprising: The computer program is executed by the processor to implement the steps in the method of any one of claims 1 to 18.
21. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps in the method of any one of claims 1 to 18.
Citation Information
Patent Citations
Video monitoring method and system for safety production
CN111325119A
Image data processing method and device and related equipment
CN111738735A
Open fire recognition method, device and equipment and storage medium
CN112598071A