A method for detecting and analyzing left objects based on video data
Through the legacy detection and analysis method based on video data, image data is collected and labeled, image segmentation and model construction are solved, and the problem of difficult to identify all legacy on high-speed road surfaces in the prior art is solved, and efficient and accurate legacy recognition is achieved.
Patent Information
- Application Number
- CN202210154772.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-21
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-02-21
AI Technical Summary
The prior art is difficult to efficiently and accurately identify all remains on high-speed road surfaces, especially non-fixed categories and various types of small-volume objects.
Using the legacy detection and analysis method based on video data, image data is collected through a preset imaging device, labeling analysis and image segmentation is carried out, and the legacy detection model is constructed, and feature extraction and recognition are performed.
It realizes efficient and accurate identification of all remains on high-speed road surfaces, improves the completeness and accuracy of identification, and enhances the breadth and acquisition efficiency of image data.
Smart Images

Figure CN114529852B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video image data detection, and particularly to a method for detecting and analyzing leftovers based on video data. Background Art
[0002] At present, deep learning technology has developed rapidly and has been widely used in the fields of image classification, natural language processing, and face recognition. Convolutional Neural Network (CNN) performs excellently in the field of computer vision and image recognition. Through the CNN network, deep image information can be extracted, and more complex high-level semantic information can be learned. For images in complex scenarios, the interference of noise can be overcome. Common network structures in convolutional neural networks include AlexNet, VGG, ResNet, GoogLeNet, etc. These networks are used to extract the features of images and finally applied to image segmentation tasks. Leftovers on the highway will cause a certain degree of interference to passing vehicles and are likely to cause traffic accidents. For leftovers in outdoor scenarios, most detection and recognition technologies are adopted. For example, in the "A Fast Leftover Detection Method and System" with the application number "201510268000.4", a filtering structure of dispersion filtering and weighted sliding average coefficient filtering is used to classify stationary targets to obtain leftovers. Since there is no fixed category for leftovers on the road surface and there are many types of leftovers, the technology of image detection and recognition can only identify fixed-category leftovers and cannot detect all leftovers. Therefore, to solve this problem, the present invention adopts the image segmentation method of deep learning to perform real-time detection and segmentation according to the leftovers on the road surface, so as to identify the leftovers. Summary of the Invention
[0003] The present invention provides a method for detecting and analyzing leftovers based on video data to solve the situation where all leftovers in the target area cannot be efficiently and accurately identified.
[0004] The present invention provides a method for detecting and analyzing leftovers based on video data, including:
[0005] Step 1: Collect data through a preset camera device to obtain image data;
[0006] Step 2: Perform annotation analysis on the image data through a preset processing system to generate an annotation strategy, and perform annotation processing to obtain an annotated image;
[0007] Step 3: Perform image segmentation on the annotated image to obtain image segmentation data;
[0008] Step 4: Construct a residue detection model based on the image segmentation data, and generate a residue detection result by extracting features from the residue detection model.
[0009] As an embodiment of this technical solution, the step 1 includes:
[0010] Deploy a preset number of camera groups within a preset range, where the camera groups include: a first camera group and a second camera group;
[0011] Through the first camera group, within the first preset time, perform data acquisition at a preset first acquisition frequency to obtain first image data;
[0012] Through the second camera group, within the second preset time, perform data acquisition at a preset second acquisition frequency to obtain second image data;
[0013] Through the first camera group and the second camera group, within the third preset time, perform data acquisition respectively at a preset third acquisition frequency, and perform data collation to generate third image data.
[0014] As an embodiment of this technical solution, the step 2 includes:
[0015] By sequentially numbering the image data, respectively obtain the corresponding numbered image groups; where;
[0016] The image data includes: first image data, second image data, and third image data;
[0017] The numbered image groups include: a first numbered image group, a second numbered image group, and a third numbered image group;
[0018] By sequentially judging the numbered images in the numbered image groups, judge whether the numbers of the numbered images are consistent with the preset numbers. If so, mark the numbered images as template images.
[0019] As an embodiment of this technical solution, the step 2 further includes:
[0020] By respectively comparing and analyzing the numbered images in the numbered image groups with the corresponding template images, obtain image deviation information, and determine an image annotation strategy according to the image deviation information; where,
[0021] The annotation strategy includes: a single annotation strategy and a multiple annotation strategy; where;
[0022] The single annotation strategy includes: performing image annotation according to a preset type of annotation;
[0023] The multi - label strategy includes: performing image labeling according to two or more preset labeling types; among them,
[0024] The labeling types include: position labeling, time labeling, and change - amount labeling.
[0025] As an embodiment of this technical solution, the step two further includes:
[0026] Through the image labeling strategy, the labeled images are sequentially subjected to labeling processing to generate labeling information;
[0027] The labeling processing includes: category labeling, position labeling, and color labeling;
[0028] According to the labeling information, labeling detection is sequentially performed to determine the labeling integrity and make a judgment; among them,
[0029] When the labeling integrity is greater than the preset threshold, the labeled image is obtained;
[0030] When the labeling integrity is less than or equal to the preset threshold, secondary labeling is performed.
[0031] As an embodiment of this technical solution, the step three includes the following steps:
[0032] Step S01: By sequentially pre - processing the labeled image, a pre - processed image is obtained; among them,
[0033] The pre - processing includes: scaling processing and de - blurring processing;
[0034] Step S02: Normalize the pre - processed image to obtain a normalized image; among them,
[0035] The normalization processing includes: traversal normalization processing and standard normal distribution processing;
[0036] Step S03: Calculate the loss of the normalized image through a preset loss function to determine the image loss value;
[0037] Step S04: According to the image loss value, classify the normalized image to determine the category region, and perform image segmentation to generate a segmented image and obtain image segmentation data.
[0038] As an embodiment of this technical solution, the image segmentation includes:
[0039] By transmitting the normalized image to a preset neural network model for image training, image training information is obtained; among them,
[0040] The image training information includes: distribution information and position information;
[0041] Construct a spatial weight matrix through a preset spatial rotation attention module and the image training information, perform feature extraction according to the spatial weight matrix, and obtain first feature data;
[0042] Perform structured analysis on the training image through a preset translation image module to obtain structured feature data, and perform feature fusion on the structured feature data and the first feature data to generate second feature data;
[0043] Perform instance segmentation according to the first feature data and the second feature data to generate a segmented image and obtain image segmentation data.
[0044] As an embodiment of this technical solution, in step four, it includes image pixel extraction:
[0045] The image pixel extraction includes the following steps:
[0046] Step S10: Compress the segmented picture according to preset conditions to obtain compressed data of the segmented picture, and determine image parameters according to the compressed data and the image segmentation data; wherein,
[0047] The compressed data includes: compressed horizontal axis data and compressed vertical axis data;
[0048] The image parameters include: image channels and image dimensions;
[0049] Step S20: Compare the image parameters with preset control parameters to obtain a parameter comparison value and make a judgment; wherein,
[0050] When the parameter comparison value is within the preset threshold range, obtain image pixel parameters;
[0051] When the parameter comparison value is not within the preset threshold range, perform improvement processing on the image parameters to obtain image pixel parameters;
[0052] Step S30: Establish a first detection model for leftovers through the image pixel parameters, and perform embedded feature extraction according to the first detection model for leftovers to obtain first image embedded features.
[0053] As an embodiment of this technical solution, step four further includes:
[0054] Screen through the image segmentation data and a preset image patch data comparison table to obtain image patch data, and input the image patch data into the embedding layer of a preset model to obtain a group of image data vectors; wherein,
[0055] The data vectors in the group of image data vectors include token vectors;
[0056] By performing superposition processing on the token vectors and sorting them, sequence images are obtained. According to the sorting order, the sequence images and a preset reference image are stacked a preset number of times to generate a superposition processed image, and feature analysis is performed to obtain the second image embedding feature.
[0057] As an embodiment of this technical solution, step four further includes:
[0058] By inputting the first image embedding feature and the second image embedding feature into a preset segmentation module and invoking a preset softmax function, image partial estimation is performed to obtain the image partial probability; wherein,
[0059] The segmentation module is a fully connected neural network structure;
[0060] By performing classification analysis on the image according to the image segmentation data, the image classification loss is obtained;
[0061] By performing transformation analysis on the image partial probability and the image classification loss through a preset multi-layer perceptron and performing prediction analysis through a preset activation function, the prediction data of the leftover image is obtained to determine the leftover recognition information.
[0062] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by practicing the present invention. The objectives and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written specification and the drawings.
[0063] The technical solution of the present invention will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0064] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:
[0065] Figure 1 is a flowchart of a leftover detection and analysis method based on video data in an embodiment of the present invention;
[0066] Figure 2 is a flowchart of step three in a leftover detection and analysis method based on video data in an embodiment of the present invention;
[0067] Figure 3 is a flowchart of image pixel extraction in a leftover detection and analysis method based on video data in an embodiment of the present invention. Detailed Embodiments
[0068] The preferred embodiments of the present invention will be described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only for the purpose of illustrating and explaining the present invention, and are not intended to limit the present invention.
[0069] It should be noted that when a component is referred to as "fixed to" or "disposed on" another component, it can be directly on the other component or indirectly on the other component. When a component is referred to as "connected to" another component, it can be directly or indirectly connected to the other component.
[0070] It should be understood that the orientation or positional relationship indicated by terms such as "length", "width", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the accompanying drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the present invention.
[0071] In addition, it should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The meaning of "a plurality" is two or more, unless otherwise specifically defined. Moreover, the terms "comprising", "including" or any other variation thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device.
[0072] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
[0073] The present invention can adapt different numbers and different position distributions of camera devices according to the acquisition scenarios of video data, obtain image data that meets the requirements of the scenarios and acquisition tasks; for different image data, generate different annotation strategies to annotate the images. While ensuring the accuracy of image data annotation, after training with a large amount of data, the annotation analysis speed and image annotation speed will become faster and faster; then perform image segmentation on the annotated images, establish a residue detection model based on the segmentation data. For different residue situations with different complexities in different scenarios, the finally generated residue detection models are also different. At this time, feature extraction is performed on the residue detection model to generate detection results, which can improve the accuracy of residue detection results in scenarios with different complexities.
[0074] Embodiment 1:
[0075] An embodiment of the present invention provides a method for detecting and analyzing residues based on video data, including:
[0076] Step 1: Collect data through a preset camera device to obtain image data;
[0077] Step 2: Perform annotation analysis on the image data through a preset processing system, generate an annotation strategy, and perform annotation processing to obtain an annotated image;
[0078] Step 3: Perform image segmentation on the annotated image to obtain image segmentation data;
[0079] Step 4: Construct a residue detection model based on the image segmentation data, perform feature extraction on the residue detection model, and generate a residue detection result;
[0080] The working principle of the above technical solution is as follows: In the prior art solution, generally after obtaining an image through video, each part of the image is recognized starting from the key points or starting from the endpoints. When encountering a complex image, the information in the image cannot be fully recognized; in the above technical solution, such as Figure 1As shown in the figure, images are collected by a camera device. Ten cameras are set up and deployed at different highway intersections to collect data. The data from 8:00 to 12:00 in the morning is collected. For highway intersections with the same perspective, one image is stored every second and defined. The data from 12:00 to 17:00 in the afternoon is collected. For highway intersections with different perspectives, one image is stored every 2 seconds and defined. And so on, all 10 cameras complete data collection and store the image data. Then, the image data is labeled and analyzed to generate a labeling strategy and perform labeling processing to obtain labeled images. The data stored under each camera is numbered. For the images numbered 1, one image is selected as the template image, and the rest of the images are used as samples for proxy labeling. According to the selected template, the position information of the objects different from the template is marked. In turn, the images of the 10 cameras are labeled. The labeled images are segmented to obtain image segmentation data. The image segmentation uniformly processes and normalizes the labeled images to obtain image segmentation data. Finally, based on the image segmentation data, a leftover detection model is constructed. By extracting features from the leftover detection model, the leftover detection results are generated.
[0081] The beneficial effects of the above technical solution are as follows: By labeling and analyzing the image data, the accuracy and efficiency of the labeling strategy are improved, and an operation basis is provided for subsequent image segmentation. By segmenting the labeled images, the image processing efficiency and the integrity of image leftover recognition are improved. By constructing a leftover detection model and extracting features, the accuracy and comprehensiveness of leftover detection are improved.
[0082] Example 2:
[0083] In one embodiment, step one includes:
[0084] By deploying a preset number of camera groups within a preset range, the camera groups include: the first camera group and the second camera group;
[0085] Through the first camera group, within the first preset time, data is collected at the preset first collection frequency to obtain the first image data;
[0086] Through the second camera group, within the second preset time, data is collected at the preset second collection frequency to obtain the second image data;
[0087] Through the first camera group and the second camera group, within the third preset time, data is collected at the preset third collection frequency respectively, and the data is sorted to generate the third image data;
[0088] The working principle of the above technical solution is as follows: In the existing technical solution, generally, multiple devices are connected to collect data simultaneously to obtain picture data and perform analysis. This method has a good recognition effect on large-volume objects in the picture, but when dealing with various small-volume objects, it cannot comprehensively recognize them. In the above technical solution, by setting up a camera group: the first camera group and the second camera group, at this time, two to ten camera groups can be set. The task of the first camera group is to collect data according to a fixed frequency within the first preset time, and the remaining camera groups respectively collect data at corresponding fixed frequencies within their corresponding times. Finally, all the data is compared and sorted to generate the final image data. In the present invention, it is the third image data. If there are ten camera groups, then the tenth image data is obtained at this time;
[0089] The beneficial effects of the above technical solution are as follows: Through the setting of multiple camera groups, with one fixed collection and the rest for auxiliary comparison collection, and finally comprehensive analysis, it greatly improves the data breadth of the recognized image, improves the integrity of the image data, and enhances the image data collection efficiency.
[0090] Embodiment 3:
[0091] In one embodiment, step two includes:
[0092] By sequentially numbering the image data, the corresponding numbered image groups are respectively obtained; among them;
[0093] The image data includes: the first image data, the second image data, and the third image data;
[0094] The numbered image groups include: the first numbered image group, the second numbered image group, and the third numbered image group;
[0095] By sequentially judging the numbered images in the numbered image groups, it is judged whether the number of the numbered image is consistent with the preset number. If so, the numbered image is marked as a template image;
[0096] The working principle of the above technical solution is as follows: In the existing technical solution, generally, images are classified to obtain image classification information. This method can quickly recognize objects of common categories, but it is very difficult to recognize, or even cannot recognize, small niche objects. In the above technical solution, the images are sequentially numbered to obtain the corresponding numbered image groups. There are multiple numbered images in the numbered image groups, and they are sequentially judged to determine whether the number of the numbered image is consistent with the preset number. If so, the numbered image is marked as a template image;
[0097] The beneficial effects of the above technical solution are as follows: By numbering the images, the usability of the images is greatly improved, and the image usage efficiency is enhanced. Through number judgment, the applicability of the images is improved, laying a good foundation for subsequent image analysis.
[0098] Example 4:
[0099] In one embodiment, step two further includes:
[0100] By comparing and analyzing the label images in the label image group with the corresponding template images respectively, obtaining image deviation information, and determining an image annotation strategy according to the image deviation information; wherein,
[0101] The annotation strategy includes: single annotation strategy, multiple annotation strategy; wherein;
[0102] The single annotation strategy includes: performing image annotation according to a preset annotation type;
[0103] The multiple annotation strategy includes: performing image annotation according to two or more preset annotation types; wherein,
[0104] The annotation types include: position annotation, time annotation, change amount annotation;
[0105] The working principle of the above technical solution is as follows: In the general technical solution, objects in the image are identified or marked sequentially from top to bottom, which has limitations and inefficiencies in use. In the above technical solution, by comparing and analyzing the label images in the label image group with the corresponding template images respectively, obtaining image deviation information, and determining an image annotation strategy according to the image deviation information, the annotation strategy includes: single annotation strategy, multiple annotation strategy, the single annotation strategy includes: performing image annotation according to a preset annotation type, the multiple annotation strategy includes: performing image annotation according to two or more preset annotation types, and the annotation types include: position annotation, time annotation, change amount annotation;
[0106] The beneficial effects of the above technical solution are: Through multiple annotation strategies, the accuracy and efficiency of image annotation are improved. Through the multiple annotation strategy, the feasibility of image annotation is improved, and the applicable range of annotation is expanded.
[0107] Example 5:
[0108] In one embodiment, step two further includes:
[0109] By means of the image annotation strategy, the label images are sequentially subjected to annotation processing to generate annotation information;
[0110] The annotation processing includes: category annotation, position annotation, color annotation;
[0111] According to the annotation information, annotation detection is sequentially performed to determine the annotation integrity and make a judgment; wherein,
[0112] When the annotation integrity is greater than a preset threshold, obtain the annotated image;
[0113] When the annotation integrity is less than or equal to the preset threshold, perform secondary annotation;
[0114] The working principle of the above technical solution is as follows: Through the image annotation strategy, the label images are sequentially annotated to generate annotation information, including: category annotation, position annotation, and color annotation; According to the annotation information, annotation detection is sequentially performed to determine the annotation integrity and make a judgment. When the annotation integrity is greater than the preset threshold, obtain the annotated image. When the annotation integrity is less than or equal to the preset threshold, perform secondary annotation; Detect the marker information in the image and at the same time judge the image annotation to determine whether it conforms to the preset annotation criteria. It should be noted here that the annotation criteria can be adjusted and modified according to the purpose of image detection; The determination of the annotation integrity also includes detecting the annotation accuracy; Among them,
[0115] The detection of the annotation accuracy includes the following steps:
[0116] Step 100: Obtain the label image group {x 1 , x 2 ,..., x n} and the preset annotation group {y 1 , y 2 ,..., y m}, and establish a system of equations:
[0117]
[0118] Among them, x i is the i-th image in the label image group, x d is the reference image, is the reference value between the i-th image in the label image group and the reference image; θ is the first reference coefficient, μ is the second reference coefficient, e is the natural logarithm, is the similarity between the i-th image in the label image group and the reference image, is the correlation value between the i-th image in the label image group and the reference image, is the variance of the first i images in the image group, is the variance of the reference image, is the covariance between the i-th image in the label image group and the reference image;
[0119] Step S200: Solve the similarity according to the system of equations
[0120]
[0121] Among them, τ is the similarity coefficient;
[0122] Step S300: According to the similarity Calculate the accuracy ω of image annotation:
[0123]
[0124] where p i is the number of correct annotations among all the annotations of the i-th image in the labeled image group, and t i is the predicted annotation quantity of the i-th image in the labeled image group, and c i is the actual annotation quantity of the i-th image in the labeled image group;
[0125] The beneficial effects of the above technical solution are as follows: Through annotation processing, the recognition rate of the image is improved, the feasibility of image processing is improved, and by judging the annotation integrity, the annotation error rate is reduced.
[0126] Example 6:
[0127] In one embodiment, step three includes the following steps:
[0128] Step S01: Obtain a preprocessed image by sequentially preprocessing the annotated image; where
[0129] The preprocessing includes: scaling processing, deblurring processing;
[0130] Step S02: Perform normalization processing on the preprocessed image to obtain a normalized image; where
[0131] The normalization processing includes: traversal standardization processing, standard normal distribution processing;
[0132] Step S03: Calculate the loss of the normalized image through a preset loss function to determine the image loss value;
[0133] Step S04: Classify the normalized image according to the image loss value to determine the class region, and perform image segmentation to generate a segmented image and obtain image segmentation data;
[0134] The working principle of the above technical solution is: As Figure 2As shown in the figure, first, the labeled image is successively subjected to scaling processing and deblurring processing to obtain a preprocessed image; second, the preprocessed image is subjected to traversal normalization processing and standard normal distribution processing to obtain a normalized image; then, the loss of the normalized image is calculated through a preset loss function to determine the image loss value; finally, according to the image loss value, the normalized image is classified to determine the category region, and image segmentation is performed to generate a segmented image and obtain image segmentation data; for the image scaling processing, the specific scaling type can be adjusted according to different images. At the same time, the loss function needs to be analyzed according to the occlusion situation of the remaining objects in the image. If the remaining objects in the image are large remaining objects, the calculation of the loss function can be appropriately adjusted or even the loss function calculation can be not performed;
[0135] The beneficial effects of the above technical solution are: by normalizing the image, the accuracy of image analysis is improved, and through the loss function, the difficulty of image analysis is reduced and the accuracy of the image is improved. Finally, by segmenting the image, the difficulty of identifying the remaining objects in the image is greatly reduced.
[0136] Example 7:
[0137] In one embodiment, the image segmentation includes:
[0138] By transmitting the normalized image to a preset neural network model for image training to obtain image training information; where
[0139] The image training information includes: distribution information, position information;
[0140] By using a preset spatial rotation attention module and the image training information, a spatial weight matrix is constructed, and feature extraction is performed according to the spatial weight matrix to obtain first feature data;
[0141] By using a preset translation image module to perform structured analysis on the training image to obtain structured feature data, and fusing the structured feature data with the first feature data to generate second feature data;
[0142] According to the first feature data and the second feature data, instance segmentation is performed to generate a segmented image and obtain image segmentation data;
[0143] The working principle of the above technical solution is as follows: Through the neural network model, image training is performed on the normalized image to obtain image distribution information and position information. Based on the preset spatial rotation attention module and the image training information, a spatial weight matrix is constructed. Feature extraction is carried out according to the spatial weight matrix to obtain the first feature data. Through the preset translation image module, structural analysis is performed on the training image to obtain structural feature data. The structural feature data is fused with the first feature data to generate the second feature data; by combining the first feature data and the second feature data, instance segmentation is carried out to generate a segmented image and obtain image segmentation data. Different colors are used to distinguish the residues with different labels in the image during instance segmentation.
[0144] The beneficial effects of the above technical solution are as follows: Through the training image, the accuracy of the distribution information and position information in the image is improved. Through the spatial rotation attention module, the efficiency of obtaining the first features of the image and the accuracy of the first features are improved. Through instance segmentation, the image annotation is made clearer and the image segmentation efficiency is improved.
[0145] Example 8:
[0146] In one embodiment, step four includes image pixel extraction:
[0147] The image pixel extraction includes the following steps:
[0148] Step S10: Compress the segmented picture according to the preset conditions to obtain the compressed data of the segmented picture. Determine the image parameters based on the compressed data and the image segmentation data. Among them,
[0149] The compressed data includes: compressed horizontal axis data and compressed vertical axis data;
[0150] The image parameters include: image channels and image dimensions;
[0151] Step S20: Compare the image parameters with the preset reference parameters to obtain a parameter comparison value and make a judgment. Among them,
[0152] When the parameter comparison value is within the preset threshold range, obtain the image pixel parameters;
[0153] When the parameter comparison value is not within the preset threshold range, then perform improvement processing on the image parameters to obtain the image pixel parameters;
[0154] Step S30: Establish a first detection model for residues based on the image pixel parameters, and perform embedded feature extraction according to the first detection model for residues to obtain the first embedded features of the image;
[0155] The working principle of the above technical solution is as follows: AsFigure 3 As shown, first, compress the segmented image, and then determine the image parameters according to the image segmentation data, including: the image channel and the image dimension; then compare the image parameters with the preset standard parameters, calculate the parameter comparison value, and make a judgment to determine whether the parameter comparison value is within the preset threshold range. If so, obtain the image pixel parameters; otherwise, perform enhancement processing, which includes enhancing the image channel and enhancing the image dimension. Finally, construct the first detection model for the residue based on the above data, and perform feature extraction to determine the first embedded feature of the image;
[0156] The beneficial effects of the above technical solution are as follows: By compressing the data, the applicability of the data is improved; by comparing the image parameters, the accuracy and matching degree of image analysis are improved; by constructing a model and extracting the embedded feature, the residue recognition efficiency is improved.
[0157] Example 9:
[0158] In one embodiment, step four further includes:
[0159] Screen through the image segmentation data and the preset image patch data comparison table to obtain the image patch data, and input the image patch data into the embedding layer of the preset model to obtain the image data vector group; where
[0160] The data vectors in the image data vector group include token vectors;
[0161] Perform superposition processing on the token vectors and sort them to obtain the sequence image. According to the sorting order, perform image stacking on the sequence image and the preset reference image a preset number of times to generate the superposition processed image, and perform feature analysis to obtain the second embedded feature of the image;
[0162] The working principle of the above technical solution is as follows: Analyze the image segmentation data according to the image patch data comparison table to determine the image patch data, and input it into the embedding layer of the preset model to obtain the image data vector group. The data vectors in the image data vector group include token vectors; then perform superposition processing on the token vectors and sort them to obtain the sequence image; finally, according to the sorting order, perform image stacking on the sequence image and the preset reference image a preset number of times to generate the superposition processed image, and perform feature analysis to obtain the second embedded feature of the image
[0163] The beneficial effects of the above technical solution are as follows: Through the image patch, the integrity of image residue recognition is improved; through the token vector, the difficulty of image processing is reduced, and the image processing speed is improved.
[0164] Example 10:
[0165] In one embodiment, step four further includes:
[0166] By inputting the first image embedding feature and the second image embedding feature into a preset segmentation module and invoking a preset softmax function, perform image part estimation to obtain the image part probability; wherein,
[0167] The segmentation module is a fully-connected neural network structure;
[0168] By performing classification analysis on the image according to the image segmentation data, obtain the image classification loss;
[0169] By performing transformation analysis on the image part probability and the image classification loss through a preset multi-layer perceptron and performing prediction analysis through a preset activation function, obtain the predicted data of the leftover image and determine the leftover identification information;
[0170] The working principle of the above technical solution is as follows: For the first image embedding feature and the second image embedding feature obtained above, perform softmax function calculation through a preset segmentation module to complete image part estimation and obtain the image part probability; the segmentation module here is a fully-connected neural network structure, then perform classification analysis on the image according to the image segmentation data to obtain the image classification loss, next perform transformation analysis on the image part probability and the image classification loss through a preset multi-layer perceptron, and finally perform prediction analysis through a preset activation function to obtain the predicted data of the leftover image and determine the leftover identification information;
[0171] The beneficial effects of the above technical solution are as follows: By processing the image through the segmentation module, the comprehensiveness of leftover image recognition is improved. Through the transformation analysis of the multi-layer perceptron and the image classification loss, the probability of deviation in leftover image calculation is reduced, and the leftover recognition efficiency is improved.
[0172] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.
[0173] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0174] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0175] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0176] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these changes and modifications.
Claims
1. A method for detecting and analyzing left-behind objects based on video data, including: Step 1: Collect data through a preset camera device to obtain image data; Step 2: Perform annotation analysis on the image data through a preset processing system, generate an annotation strategy, and perform annotation processing to obtain an annotated image; Step 3: Perform image segmentation on the annotated image to obtain image segmentation data; Step 4: According to the image segmentation data, construct a left-behind object detection model, and generate a left-behind object detection result by performing feature extraction on the left-behind object detection model; The said Step 1 includes: Deploy a preset number of camera groups within a preset range, and the camera groups include: the first camera group and the second camera group; Through the first camera group, within the first preset time, collect data at a preset first collection frequency to obtain first image data; Through the second camera group, within the second preset time, collect data at a preset second collection frequency to obtain second image data; Through the first camera group and the second camera group, within the third preset time, collect data respectively at a preset third collection frequency, and perform data sorting to generate third image data; The said Step 2 further includes: Perform annotation processing on the numbered image in sequence through an image annotation strategy to generate annotation information; The said annotation processing includes: category annotation, position annotation, color annotation; Perform annotation detection in sequence according to the annotation information, determine the annotation integrity, and make a judgment; wherein, When the annotation integrity is greater than a preset threshold, then obtain the annotated image; When the annotation integrity is less than or equal to the preset threshold, then perform secondary annotation; The said Step 3 includes the following steps: Step S01: Perform preprocessing on the annotated image in sequence to obtain a preprocessed image; wherein, The said preprocessing includes: scaling processing, deblurring processing; Step S02: Perform normalization processing on the preprocessed image to obtain a normalized image; wherein, The said normalization processing includes: traversal standardization processing, standard normal distribution processing; Step S03: Calculate the loss of the normalized image through a preset loss function to determine the image loss value; Step S04: Perform category division on the normalized image according to the image loss value, determine the category area, and perform image segmentation to generate a segmented image, and obtain image segmentation data.
2. A method for detecting and analyzing left-behind objects based on video data as described in claim 1, characterized in that the said Step 2 includes: Number the image data in sequence to obtain corresponding numbered image groups respectively; wherein; The said image data includes: first image data, second image data, third image data; The said numbered image groups include: first numbered image group, second numbered image group, third numbered image group; Judge the numbers of the numbered images in the numbered image groups in sequence to judge whether the numbers of the numbered images are consistent with the preset numbers. If so, mark the numbered images as template images.
3. A method for detecting and analyzing left-behind objects based on video data as described in claim 1, characterized in that the said Step 2 further includes: By comparing and analyzing the labeled images in the labeled image group with the corresponding template images respectively, image deviation information is obtained, and according to the image deviation information, an image annotation strategy is determined; wherein, The annotation strategies include: single annotation strategy, multiple annotation strategy; wherein; The single annotation strategy includes: performing image annotation according to a preset annotation type. The multiple annotation strategy includes: performing image annotation according to two or more preset annotation types; wherein, The annotation types include: position annotation, time annotation, change amount annotation.
4. The method for detecting and analyzing leftovers based on video data according to claim 1, characterized in that the image segmentation includes: By transmitting the normalized image to a preset neural network model for image training, image training information is obtained; wherein, The image training information includes: distribution information, position information; By means of a preset spatial rotation attention module and the image training information, a spatial weight matrix is constructed, and feature extraction is performed according to the spatial weight matrix to obtain first feature data; By means of a preset translation image module, a structured analysis of the training image is performed to obtain structured feature data, and the structured feature data is fused with the first feature data to generate second feature data; According to the first feature data and the second feature data, instance segmentation is performed to generate a segmented image, and image segmentation data is obtained.
5. The method for detecting and analyzing leftovers based on video data according to claim 1, characterized in that Step four includes image pixel extraction: The image pixel extraction includes the following steps: Step S10: Compress the segmented picture according to preset conditions to obtain compressed data of the segmented picture, and determine image parameters according to the compressed data and the image segmentation data; wherein, The compressed data includes: compressed horizontal axis data, compressed vertical axis data; The image parameters include: image channels, image dimensions; Step S20: Compare the image parameters with preset reference parameters to obtain a parameter comparison value and make a judgment; wherein, When the parameter comparison value is within a preset threshold range, image pixel parameters are obtained; When the parameter comparison value is not within the preset threshold range, the image parameters are enhanced to obtain image pixel parameters; Step S30: Establish a first detection model for leftovers through the image pixel parameters, and perform embedded feature extraction according to the first detection model for leftovers to obtain first image embedded features.
6. The method for detecting and analyzing leftovers based on video data according to claim 1, characterized in that Step four further includes: By screening through the image segmentation data and a preset image patch data comparison table, image patch data is obtained, and the image patch data is input into the embedding layer of a preset model to obtain an image data vector group; wherein, The data vectors in the image data vector group include token vectors. By performing superposition processing on the token vectors and sorting them, sequence images are obtained. According to the sorting order, the sequence images and a preset reference image are stacked a preset number of times to generate a superposition processed image, and feature analysis is performed to obtain the second image embedding feature.
7. A method for detecting and analyzing left-behind objects based on video data according to claim 1, wherein, the fourth step further includes: By inputting the first image embedding feature and the second image embedding feature into a preset segmentation module and invoking a preset softmax function, image partial estimation is performed to obtain the image partial probability; wherein, the segmentation module is a fully connected neural network structure; By performing classification analysis on the image according to the image segmentation data, the image classification loss is obtained; By performing conversion analysis on the image partial probability and the image classification loss through a preset multi-layer perceptron and performing prediction analysis through a preset activation function, the left-behind object image prediction data is obtained to determine the left-behind object recognition information.
Citation Information
Patent Citations
A fast method and system for detecting residues
CN104881643B
Left-behind object detection method and device and equipment
CN112837326A
Method and device for detecting left article based on fusion algorithm
CN113393482A
Method and apparatus for target object segmentation in image, and electronic device and storage medium
WO2021189913A1