Artificial Intelligence-Based Optimized Acquisition Method for Factory Videos

By setting the loss tolerance of sparse processing in a smart factory according to the monitoring location and time period, and sparse the monitoring video frames multiple times, the problems of large data volume and high bandwidth pressure during the monitoring video transmission process are solved, real-time transmission of video and clear identification of abnormal situations are achieved, and the reliability and efficiency of monitoring are improved.

CN120050397BActive Publication Date: 2025-07-22XIAN XINGXUN INTELLIGENT COMM TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510512100.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-22
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

In a smart factory, during the acquisition and transmission of surveillance videos, how to balance the video quality and compression level to solve the problem of insufficient real-time due to large data volume and high bandwidth pressure, and ensure that surveillance videos can be transmitted in a timely manner and that surveillance personnel can clearly identify abnormal situations.

Method used

Install surveillance cameras at different monitoring locations in the factory, set the loss tolerance of sparse processing according to the monitoring location and the monitoring level of the video acquisition period, and sparse the monitoring video frames multiple times until the loss level is less than the tolerance level, obtain the optimized acquisition results, and transmit supplementary information to support the accurate restoration of the monitoring center.

Benefits of technology

Lossy compression of surveillance video is realized, real-time and reliability of video transmission is ensured, data volume is reduced, efficiency and security of surveillance video is improved, and efficient operation and security management of smart factories are supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050397B_ABST
    Figure CN120050397B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of image processing, and specifically relates to an optimized acquisition method for factory videos based on artificial intelligence. The method includes: installing monitoring cameras at different monitoring positions in the factory, setting the loss tolerance δ for sparse processing according to the monitoring positions and the monitoring levels during the acquisition period of the monitoring videos; performing multiple sparse processes on the initial video frames in the monitoring videos: dividing all pixel points in the video frame after the previous sparse process into two categories, and the 4 neighboring pixel points of any pixel point are different from the category to which the pixel point belongs, and forming the video frame after the current sparse process with the category having the minimum loss degree until the loss degree of the video frame after the R-th sparse process is less than δ and the loss degree of the video frame after the (R + 1)-th sparse process is not less than δ, taking the video frame after the R-th sparse process as the optimized acquisition result and transmitting it. The present invention achieves a balance between video quality and compression degree.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing. More specifically, the present invention relates to an optimized acquisition method for factory videos based on artificial intelligence. Background Art

[0002] A smart factory refers to an intelligent factory that establishes a virtual model and a simulation environment based on digital twin technology to optimize factory production and operation; through the digital twin smart factory, manufacturing enterprises can simulate the entire production process in a virtual environment, including equipment operation, material flow, personnel management, etc. This enables enterprises to better predict potential problems, optimize production processes, improve efficiency and quality, and at the same time reduce production costs.

[0003] In the construction of a smart factory, the acquisition of surveillance videos plays a crucial role in remote security monitoring. Through the acquisition and analysis of surveillance videos, a smart factory can achieve visual and intelligent management of the production process, improve production efficiency, safety, and quality, reduce operating costs, and enhance the overall competitiveness of the enterprise; at the same time, the acquisition and transmission of surveillance videos are the key links to achieve visual and intelligent management of the production process.

[0004] However, in practical applications, due to the large amount of surveillance video data, simultaneous acquisition at multiple locations will cause high network bandwidth pressure, resulting in transmission delays, affecting real-time performance, and further leading to missed opportunities for handling critical events; therefore, it is necessary to compress and transmit the surveillance videos.

[0005] When compressing surveillance videos, excessive compression will reduce the quality of the surveillance videos. Insufficient quality of the surveillance videos will prevent surveillance personnel from clearly identifying abnormal situations, affecting the reliability of surveillance; while insufficient compression will result in too large a data volume, increasing the transmission burden, affecting real-time performance, and further leading to missed opportunities for handling critical events. Summary of the Invention

[0006] To solve the above technical problem of how to balance video quality and compression degree, the present invention provides an optimized acquisition method for factory videos based on artificial intelligence, including: installing surveillance cameras at different monitoring positions in the factory, and setting the loss tolerance of sparse processing according to the monitoring level of the monitoring position and the monitoring level of the acquisition period of the surveillance video ; performing multiple sparse processes on the video frames in the surveillance videos collected by the surveillance cameras until and and , taking the video frame as the optimized acquisition result and transmitting it, is the loss degree of the video frame after the th sparse process; the th sparse process includes: taking the The video frames after sub-sparse processing All pixel points in it are divided into two categories, and it is required that all neighboring pixel points within the 4-neighborhood of any pixel point are different from the category to which the pixel point belongs; calculate the loss degree when retaining any category; form a video frame with all pixel points in the category with the minimum loss degree and use the minimum loss degree as the loss degree of the video frame Calculating the loss degree includes: calculating the contribution coefficient of each neighborhood position within the 4-neighborhood during the th sub-sparse processing according to the gray-scale difference between all pixel points in the non-target category of the two categories and each neighboring pixel point within their 4-neighborhoods; according to the contribution coefficients of each neighborhood position within the 4-neighborhood during the previous th sub-sparse processing and the gray-scale values of all pixel points in the target category, obtain a restored frame of the same size as the video frame through multiple restorations, and calculate the loss degree when retaining the target category according to the difference in the gray-scale values of the pixel points in the restored frame and the video frame .

[0007] In the present invention, by performing multiple sub-sparse processings on the collected video frames, obtaining an optimized acquisition result of the video frames and transmitting them, lossy compression of the video frames is achieved, effectively solving the problem of insufficient real-time performance caused by large amounts of data and high bandwidth pressure during the acquisition and transmission of monitoring videos in smart factories, ensuring that the monitoring videos can be transmitted to the monitoring center in real time, enabling monitoring personnel to promptly discover and handle abnormal situations during the production process, and improving production efficiency and safety; during this process, according to the monitoring level of the monitoring location and the monitoring level of the acquisition period of the monitoring video, the loss tolerance of the sub-sparse processing is dynamically adjusted until a video frame with the maximum loss degree of the sub-sparse processed video frame and less than the loss tolerance is obtained, ensuring that the video quality under different monitoring requirements is within an acceptable range, reducing both the amount of data to be transmitted and retaining the key information in the monitoring videos, enabling monitoring personnel to clearly identify abnormal situations and ensuring the reliability of monitoring.

[0008] Preferably, the loss tolerance of the sub-sparse processing , is the monitoring degree corresponding to the monitoring level of the monitoring location, is the monitoring degree corresponding to the monitoring level of the acquisition period of the monitoring video, is the upper limit of the loss tolerance; the monitoring level includes three levels, namely high level, medium level, and low level, and the monitoring degrees corresponding to the three monitoring levels are respectively , , , and .

[0009] The present invention dynamically adjusts the loss tolerance of sparse processing according to the monitoring level of the monitoring location and the monitoring level of the acquisition period, and can ensure that the video quality meets the requirements under different monitoring needs.

[0010] Preferably, all pixel points in the video frame after the th sparse processing are divided into two categories, and it is required that all neighboring pixel points within the 4-neighborhood of any pixel point are different from the category to which the pixel point belongs, including: for all pixel points in the video frame , the pixel points whose row and column are odd rows and odd columns and the pixel points whose row and column are even rows and even columns are divided into the first category; the pixel points whose row and column are odd rows and even columns and the pixel points whose row and column are even rows and odd columns are divided into the second category.

[0011] Preferably, calculating the contribution coefficient of each neighboring position within the 4-neighborhood during the th sparse processing includes: the contribution coefficient of the neighboring position within the 4-neighborhood during the th sparse processing ; where is the average gray difference of the neighboring pixel points at the neighboring position within the 4-neighborhood during the th sparse processing, , is equal to , , , and the sum of

[0012] The present invention calculates the gray difference between a pixel point and each neighboring pixel point within its 4-neighborhood, and calculates the contribution coefficient of each neighboring position within the 4-neighborhood during sparse processing, so as to ensure that when the video frame is restored according to the contribution coefficient of each neighboring position within the 4-neighborhood during sparse processing and the gray values of the remaining pixel points, the difference between the restored frame and the original video frame is small, ensuring the video quality.

[0013] Preferably, obtaining a restored frame of the same size as the video frame through multiple restorations includes: setting a blank image equal in size to the video frame , from to 1; according to the positions of the remaining pixel points in the video frame , setting all pixel points in the restored frame in the image to obtain the image to be restored ; through the When performing sub-sparse processing, the contribution coefficients of each neighborhood position in the 4-neighborhood are used to perform weighted summation on the gray values of all neighborhood pixels of each blank pixel in the image to be restored, so as to obtain the gray value of each blank pixel in the image to be restored, and then obtain the gray value of each blank pixel in the restored frame; until is equal to 1 and then stop, to obtain the restored frame , and the restored frame has the same size as the video frame . Preferably, obtaining the gray value of each blank pixel in the image to be restored includes: ; in the formula,

[0014] is the gray value of the blank pixel, is the number of neighborhood positions where neighborhood pixels exist in the 4-neighborhood of the blank pixel, is the contribution coefficient of the th neighborhood position where neighborhood pixels exist in the 4-neighborhood of the blank pixel, is the gray value of the th neighborhood pixel existing in the 4-neighborhood of the blank pixel, is the contribution coefficient of the th neighborhood position where no neighborhood pixel exists in the 4-neighborhood of the blank pixel, is the average value of the gray values of all neighborhood pixels existing in the 4-neighborhood of the blank pixel,

[0015]

[0016]

[0017] Preferably, the positions of the pixels retained in the video frame

[0016] are determined according to the retention class in the video frame: when the retention class is the first class, the positions of the pixels retained are those where the row and column are odd rows and odd columns and those where the row and column are even rows and even columns; when the retention class is the second class, the positions of the pixels retained are those where the row and column are odd rows and even columns and those where the row and column are even rows and odd columns. Preferably, calculating the loss degree when retaining the target class according to the difference in the gray values of the pixels in the restored frame and the video frame includes: taking the average value of the differences in the gray values of all pixels in the restored frame and all pixels in the video frame

[0017] Preferably, the method further includes: the number of times of sparse processing during the process of obtaining the optimized acquisition result , the retained classes in the video frames after each sparse processing, and the contribution coefficients of each neighborhood position in the 4-neighborhood during each sparse processing are used as supplementary information and transmitted.

[0018] The present invention transmits the supplementary information so that the monitoring center can accurately restore according to the supplementary information and the optimized acquisition result, ensuring the reliability of monitoring.

[0019] Preferably, the method further includes: at the monitoring center, according to the optimized acquisition result and the supplementary information, obtaining the video frames for display through multiple restorations, including: setting a blank image with a size equal to ; , as the intermediate image , where the intermediate image refers to the optimized acquisition result; according to the retained classes in the video frames after the th sparse processing, determining the positions of the retained pixel points in the video frames after the th sparse processing, for setting the pixel points in the intermediate image in the blank image ; obtaining the gray values of the remaining blank pixel points by weighted summing the gray values of all the neighborhood pixel points of the remaining blank pixel points in through the contribution coefficients of each neighborhood position in the 4-neighborhood during the th sparse processing, obtaining the gray values of the remaining blank pixel points, and obtaining the nd intermediate image ; Taking from to 1 until obtaining the intermediate image , as the video frames for display.

[0020] The beneficial effects of the present invention are as follows:

[0021] By performing multiple sparse processings on the acquired video frames, the present invention obtains the optimized acquisition result of the video frames and transmits them, realizing the lossy compression of the video frames; at the same time, according to the monitoring level of the monitoring location and the monitoring level of the acquisition period of the monitoring video, dynamically adjusting the loss tolerance of the sparse processing until obtaining the video frames with the maximum loss degree of the sparse processed video frames and less than the loss tolerance, ensuring that the video quality under different monitoring requirements is within an acceptable range; therefore, the present invention not only improves the efficiency of monitoring video transmission, but also ensures the real-time performance and reliability of monitoring, providing strong support for the efficient operation and safety management of intelligent factories. Description of the Drawings

[0022] Figure 1is a flowchart schematically showing the method for optimizing the acquisition of factory videos based on artificial intelligence in the present invention;

[0023] Figure 2 is a schematic diagram schematically showing the classification of pixel points;

[0024] Figure 3 is a schematic diagram schematically showing a blank image;

[0025] Figure 4 is a schematic diagram schematically showing an image to be restored;

[0026] Figure 5 is a schematic diagram schematically showing a restored frame. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present invention.

[0028] Next, the detailed implementation manners of the present invention will be described in detail in conjunction with the accompanying drawings.

[0029] The embodiments of the present invention disclose a method for optimizing the acquisition of factory videos based on artificial intelligence. Referring to Figure 1 , it includes steps S1 to S3:

[0030] S1. Install monitoring cameras at different monitoring positions in the factory to collect monitoring videos of each monitoring position; set the loss tolerance of sparse processing according to the monitoring level of the monitoring position and the monitoring level of the acquisition period of the monitoring video.

[0031] A smart factory refers to an intelligent factory that builds a virtual model and a simulation environment based on digital twin technology to optimize factory production and operation; through the digital twin smart factory, manufacturing enterprises can simulate the entire production process in a virtual environment, including equipment operation, material flow, personnel management, etc. enabling enterprises to better predict potential problems, optimize production processes, improve efficiency and quality, and at the same time reduce production costs.

[0032] In the construction of an intelligent factory, the collection of surveillance videos plays a crucial role in remote security monitoring. By transmitting the surveillance videos to remote terminals via the network, managers can view the real-time images of the factory at any time and place through devices such as mobile phones and computers, achieving remote monitoring and management, and supporting collaborative work among multiple departments, improving work efficiency and collaborative effectiveness. At the same time, the surveillance videos can monitor the personnel activities and environmental conditions in the factory in real time, ensure that employees comply with safety operation procedures, promptly detect and handle unsafe behaviors, and guarantee the safety of personnel and equipment.

[0033] In a factory, setting multiple monitoring locations is the key to achieving comprehensive monitoring and optimizing the production process. Common monitoring locations include: production equipment areas, material storage areas, personnel work areas, safety critical areas, energy management areas, and environmental monitoring areas, and the monitoring levels are respectively: high, medium, low, high, medium, medium.

[0034] The collection period of surveillance videos can be divided according to the production plan and management requirements of the factory: usually divided into production periods and non-production periods. The production period includes 8:00 am to 5:00 pm on weekdays, and the monitoring level is high; the non-production period includes 5:00 pm on weekdays to 8:00 am the next day, and the whole day on weekends, and the monitoring level is medium.

[0035] The monitoring levels include three levels, namely high, medium, and low. The monitoring degrees corresponding to the three monitoring levels are respectively 、 、 ,and ; 、 、 The specific values of can be set according to the actual application scenarios and requirements, and the value range is [0, 1]. In this embodiment, 、 、 are respectively set to 0.95, 0.75, and 0.5.

[0036] Specifically, install surveillance cameras at different monitoring locations in the factory to collect surveillance videos of each monitoring location.

[0037] For the surveillance videos collected by the surveillance cameras, take any video frame in the surveillance video as the initial video frame, and denote it as video frame ; for video frame , taking the pixel point in the lower left corner of the video frame as the origin, the horizontal right direction from the origin as the positive x-axis direction, and the vertical upward direction from the origin as the positive y-axis direction, construct a rectangular coordinate system, and denote the coordinates of the pixel points in the video frame in the rectangular coordinate system as , is the abscissa of the pixel point, is the ordinate of the pixel point.

[0038] Furthermore, according to the monitoring level of the monitoring location and the monitoring level of the acquisition period of the monitoring video, set the loss tolerance of the sparse processing; then the loss tolerance of the sparse processing , is the monitoring degree corresponding to the monitoring level of the monitoring location, is the monitoring degree corresponding to the monitoring level of the acquisition period of the monitoring video, is the upper limit of the loss tolerance.

[0039] Among them, the specific value of the upper limit of the loss tolerance can be set according to the actual application scenario and requirements, and the value range is [40, 60]. In this embodiment, the upper limit of the loss tolerance is set to 50.

[0040] It should be noted that according to the monitoring level of the monitoring location and the monitoring level of the acquisition period of the monitoring video, dynamically adjust the loss tolerance of the sparse processing to ensure the video quality under different monitoring requirements: for high-monitoring-level areas, a lower loss tolerance can be set to ensure the video quality; for low-monitoring-level areas, a higher loss tolerance can be set to further reduce the data volume.

[0041] S2. Perform multiple sparse processings on the initial video frame in the monitoring video until the loss degree of the video frame after the current sparse processing is less than the loss tolerance and the loss degree of the video frame after the next sparse processing is not less than the loss tolerance, then use the video frame after the current sparse processing as the optimized acquisition result and transmit it.

[0042] It should be noted that by performing multiple sparse processings on the video frame , the number of pixel points in the video frame after each sparse processing will be reduced by half, thereby reducing the size of the video frame to be transmitted. Therefore, in this embodiment, through sparse processing, lossy compression of the video frame is achieved, which can significantly reduce the data volume and reduce the demand for network bandwidth, enabling the video to be transmitted more smoothly, reducing latency and stuttering, ensuring that the monitoring video can be transmitted to the monitoring center in real time, and further ensuring that the monitoring personnel can promptly discover and handle abnormal situations in the production process, improving production efficiency and safety.

[0043] Specifically, for the video frame , perform multiple sparse processings on the video frame until the th sparse processing of the video frame has a loss degree and the th sparse processing of the video frame Degree of loss , the video frame after the th sparsification process is used as the optimized acquisition result. , are respectively the video frame after the th sparsification process and the video frame after the th sparsification process Degree of loss.

[0044] It should be noted that in this embodiment, for the optimized acquisition result of the monitored video, the video frame with the largest degree of loss and less than the loss tolerance is used to ensure that the video quality under different monitoring requirements is within an acceptable range, which not only reduces the amount of data to be transmitted but also retains the key information in the monitored video, enabling the monitoring personnel to clearly identify abnormal situations and ensuring the reliability of monitoring.

[0045] Furthermore, during the process of obtaining the optimized acquisition result, the number of sparsification processes , the retained classes in the video frame after each sparsification process, and the contribution coefficients of each neighborhood position within the 4-neighborhood during each sparsification process are used as supplementary information.

[0046] In addition, the optimized acquisition result and the supplementary information are transmitted to the monitoring center.

[0047] Among them, the steps for obtaining the video frame after the th sparsification process are as follows:

[0048] 1. Denote the video frame after the th sparsification process as , divide all the pixel points in the video frame into two categories, and denote them as the first category and the second category in the video frame respectively, requiring that all the neighborhood pixel points within the 4-neighborhood of any pixel point are different from the class to which the pixel point belongs.

[0049] Specifically, for all the pixel points in the video frame , the pixel points whose row and column numbers are odd rows and odd columns and the pixel points whose row and column numbers are even rows and even columns are divided into the first category in the video frame ; the pixel points whose row and column numbers are odd rows and even columns and the pixel points whose row and column numbers are even rows and odd columns are divided into the second category in the video frame .

[0050] Exemplarily, a schematic diagram of pixel point classification is shown in Figure 2As shown, among them, the pixel points in the first category are marked as 1, and the pixel points in the second category are marked as 2.

[0051] Among them, the odd rows / odd columns refer to the rows / columns with odd serial numbers; the even rows / even columns refer to the rows / columns with even serial numbers.

[0052] For the pixel point with coordinates , its 4-neighborhood consists of 4 neighboring pixel points located above, below, left, and right of it. The 4 neighboring pixel points above, below, left, and right of the pixel point refer to the pixel points with coordinates , , , ; at the same time, the 4 positions above, below, left, and right in the 4-neighborhood are denoted as the 4 neighborhood positions in the 4-neighborhood.

[0053] 2. Calculate the loss degree when retaining any one category in the video frame .

[0054] Specifically, after all the pixel points in the video frame are divided into two categories, any one category in the video frame is taken as the target category, then the other category in the video frame is taken as the non-target category; according to the gray-scale differences between all the pixel points in the non-target category in the video frame and their respective neighboring pixel points in the 4-neighborhood, calculate the contribution coefficients of each neighborhood position in the 4-neighborhood during the th sparse processing; according to the contribution coefficients of each neighborhood position in the 4-neighborhood during the previous sparse processings and the gray-scale values of all the pixel points in the target category, obtain a restored frame with the same size as the video frame through multiple restorations; according to the differences between the pixel points in the restored frame and the gray-scale values of the pixel points in the video frame , calculate the loss degree when retaining the target category.

[0055] The specific calculation process of the loss degree when retaining the target category is as follows:

[0056] 2.1. According to the gray-scale differences between all the pixel points in the non-target category in the video frame and their respective neighboring pixel points in the 4-neighborhood, calculate the contribution coefficients of each neighborhood position in the 4-neighborhood during the th sparse processing.

[0057] For the 4-neighborhood, the 4 neighboring pixel points in the 4-neighborhood are respectively denoted as neighboring pixel point , neighboring pixel point , neighboring pixel point , neighboring pixel point ; Denote the 4 neighborhood positions within the 4-neighborhood as neighborhood position , neighborhood position , neighborhood position , neighborhood position ; And the neighborhood pixels , neighborhood pixels , neighborhood pixels , neighborhood pixels are located at neighborhood positions , neighborhood position , neighborhood position , neighborhood position .

[0058] Specifically, take any one pixel among all the pixels in the non-target class in the video frame as the concerned pixel; calculate the gray-scale difference between the concerned pixel and each neighborhood pixel within its 4-neighborhood, where the gray-scale difference refers to the absolute value of the difference in gray-scale values; thus, calculate the gray-scale differences between each pixel in the non-target class and each neighborhood pixel within its 4-neighborhood.

[0059] Furthermore, according to the gray-scale differences between all the pixels in the non-target class in the video frame and different neighborhood pixels within their 4-neighborhoods, calculate the contribution coefficients of different neighborhood positions within the 4-neighborhood at the th sparse processing, then the contribution coefficient of the neighborhood position within the 4-neighborhood at the th sparse processing; in the formula, is the average gray-scale difference of the neighborhood pixels at the neighborhood position within the 4-neighborhood at the th sparse processing, , is the average gray-scale difference of the neighborhood pixels at the neighborhood position within the 4-neighborhood at the th sparse processing, ; represents , , , sum.

[0060] Among them, the average gray-scale difference of the neighborhood pixels at the neighborhood position within the 4-neighborhood at the th sparse processing; in the formula, is the number of all the pixels in the non-target class in the video frame , is the The gray - level difference between a pixel and its neighboring pixels is described as follows.

[0061] It should be noted that, for the neighboring positions within the 4 - neighborhood the average value of the gray - level differences of the neighboring pixels is larger, the contribution coefficient of the neighboring position within the 4 - neighborhood is smaller. This ensures that when the video frame is restored based on the contribution coefficients of each neighboring position within the 4 - neighborhood and the gray - level values of the retained pixels during subsequent sparse processing, the difference between the restored frame and the original video frame is small, thus ensuring the video quality.

[0062] 2.2. Based on the contribution coefficients of each neighboring position within the 4 - neighborhood during the previous sparse processing and the gray - level values of all pixels in the target class, a restored frame with the same size as the video frame is obtained through multiple restorations.

[0063] The specific process of obtaining the restored frame through multiple restorations is as follows:

[0064] (1) Set a blank image with the same size as the video frame , as shown in Figure 3.

[0065] (2) According to the positions of the retained pixels in the video frame , fill all the pixels in the target class into the image to obtain the image to be restored , as shown in Figure 4 .

[0066] Since the image to be restored has the same size as the video frame , and the number of all pixels in the target class is equal to half of the number of all pixels in the video frame , therefore, when filling all the pixels in the target class into the image , there are still half of the pixels in the obtained image to be restored that are blank pixels, and the neighboring pixels within the 4 - neighborhood of each blank pixel are pixels in the target class.

[0067] In addition, since all the pixels in the video frame in this embodiment are divided into two classes, and it is required that all the neighboring pixels within the 4 - neighborhood of any pixel are different from the class to which the pixel belongs, therefore, for any pixel in the non - target class in the video frame , the neighboring pixels within its 4 - neighborhood are essentially pixels in the target class in the video frame .

[0068] (3) By the contribution coefficients of each neighborhood position in the 4-neighborhood during the th sparse processing, the gray values of all neighborhood pixels in the 4-neighborhood of each blank pixel in the image to be restored are weighted and summed to obtain the gray value of each blank pixel in the image to be restored. In this way, the gray value of each blank pixel in the image to be restored is obtained, and thus the th restored frame is obtained. The schematic diagram is as shown in . Figure 5

[0069] (4) Set a blank image with the same size as the video frame , where ranges from to 1.

[0070] (5) According to the positions of the retained pixels in the video frame , all the pixels in the restored frame are set in the image to obtain the image to be restored .

[0071] Since the restored frame and the video frame have the same size, the image to be restored has the same size as the video frame . And the size of the video frame is half of the size of the video frame . Therefore, there are still half of the pixels in the image to be restored that are blank pixels, and the neighborhood pixels in the 4-neighborhood of each blank pixel are the pixels in the restored frame .

[0072] (6) By the contribution coefficients of each neighborhood position in the 4-neighborhood during the th sparse processing, the gray values of all neighborhood pixels of each blank pixel in the image to be restored are weighted and summed to obtain the gray value of each blank pixel in the image to be restored. In this way, the gray value of each blank pixel in the image to be restored is obtained, and thus the th restored frame is obtained.

[0073] (7) And so on, until is equal to 1, the restored frame is obtained. The restored frame has the same size as the video frame .

[0074] Among them, the positions of the pixels retained in the video frame are determined according to the retention classes in the video frame: when the retention class is the first class, the positions of the pixels to be retained refer to the positions where the row and column are odd rows and odd columns and the positions where the row and column are even rows and even columns; when the retention class is the second class, the positions of the pixels to be retained refer to the positions where the row and column are odd rows and even columns and the positions where the row and column are even rows and odd columns.

[0075] Among them, through the contribution coefficients of each neighborhood position in the 4-neighborhood during the th sparse processing, the gray values of all neighborhood pixels of each blank pixel in the image to be restored are weighted and summed to obtain the gray value of each blank pixel in the image to be restored

[0076] ;

[0077] In the formula, is the gray value of the blank pixel, is the number of neighborhood positions where neighborhood pixels exist in the 4-neighborhood of the blank pixel, then represents the number of neighborhood positions where no neighborhood pixels exist in the 4-neighborhood of the blank pixel, is the th neighborhood position where neighborhood pixels exist in the 4-neighborhood of the blank pixel 's contribution coefficient, is the gray value of the th neighborhood pixel existing in the 4-neighborhood of the blank pixel, is the th neighborhood position where no neighborhood pixels exist in the 4-neighborhood of the blank pixel 's contribution coefficient, is the average value of the gray values of all neighborhood pixels existing in the 4-neighborhood of the blank pixel, represents rounding.

[0078] Among them, when all neighborhood positions in the 4-neighborhood of the blank pixel have neighborhood pixels, the number of neighborhood positions where neighborhood pixels exist in the 4-neighborhood of the blank pixel = 4, then the number of neighborhood positions where no neighborhood pixels exist in the 4-neighborhood of the blank pixel = 0. At this time, the restored value

[0079] It should be specifically noted that for the pixel points located in the first row / first column / last row / last column of the video frame, the number of neighborhood positions with neighborhood pixel points within their 4-neighborhoods is less than 4. That is to say, there are multiple neighborhood positions within their 4-neighborhoods that do not include neighborhood pixel points. For the neighborhood positions without neighborhood pixel points, in this embodiment, the weighted sum is obtained by averaging the gray values of all the neighborhood pixel points existing within the 4-neighborhood of the blank pixel point, so as to obtain the restored value of the blank pixel point.

[0080] 2.3. According to the restored frame and the pixel points in the video frame calculate the loss degree when retaining the target class.

[0081] Specifically, the average value of the differences in gray values between all the pixel points in the restored frame and all the pixel points in the video frame is used as the loss degree when retaining the target class.

[0082] It should be noted that the greater the difference between the gray value and the restored value, the greater the loss degree when retaining the target class.

[0083] Thus, after dividing all the pixel points in the video frame into two categories, calculate the loss degree when retaining each category.

[0084] 3. Take the category with the minimum loss degree as the retained category in the video frame ; Denote all the pixel points in the retained category in the video frame as the retained pixel points in the video frame ; And form the video frame after the -th sparse processing by all the retained pixel points in the video frame ; Take the minimum loss degree, that is, the loss degree of the retained category in the video frame as the loss degree of the video frame after the -th sparse processing.

[0085] It should be specifically noted that when , the steps to obtain the video frame after the -th sparse processing are as follows: Divide all the pixel points in the initial video frame into two categories, requiring that all the neighborhood pixel points within the 4-neighborhood of any pixel point are different from the category to which the pixel point belongs; Take any category as the target class, then the other category is taken as the non-target class; Calculate the loss degree when retaining the target class; Form the video frame after the -th sparse processing by all the pixel points in the category with the minimum loss degree, and take the minimum loss degree as the loss degree of the video frame after the -th sparse processing.

[0086] S3. At the monitoring center, based on the optimized acquisition results and supplementary information, obtain the video frames for display through multiple restorations to achieve remote security monitoring.

[0087] The process of obtaining the video frames for display through multiple restorations is as follows:

[0088] (1) Set a blank image with a size equal to ; where, ranges from to taking 1, is the number of times of sparse processing, is the size of the intermediate image ; when = , the intermediate image refers to the optimized acquisition results.

[0089] (2) Determine the positions of the retained pixel points in the video frame after the -th sparse processing according to the retained classes in the video frame after the -th sparse processing; according to the positions of the retained pixel points in the video frame after the -th sparse processing, set all the pixel points in the -th intermediate image in the blank image ; through the contribution coefficients of each neighborhood position in the 4-neighborhood during the -th sparse processing, weighted sum the gray values of all the neighborhood pixel points of the remaining blank pixel points in the blank image to obtain the gray values of the remaining blank pixel points, and thus obtain the -th intermediate image .

[0090] (3) And so on, until is equal to 1, to obtain the intermediate image , as the video frame for display.

[0091] It should be noted that according to the optimized acquisition results and supplementary information, the present invention obtains the video frames for display through multiple restorations to realize the visualization and intelligent management of the production process. Managers can view the real-time pictures of the factory through devices such as mobile phones and computers at any time and place, realizing remote monitoring and management, which is helpful for the collaborative work among multiple departments and improves work efficiency and collaborative effect.

Claims

1. An artificial intelligence-based method for optimizing the acquisition of factory videos, characterized in that, including: Install surveillance cameras at different monitoring locations in the factory, and set the loss tolerance of sparse processing according to the monitoring level of the monitoring location and the monitoring level of the acquisition period of the surveillance video ; Video frames in the surveillance video collected by the surveillance camera are sparsely processed multiple times until and , and the video frames are used as the optimized acquisition result and transmitted, is the loss degree of the video frame after the th sparse processing; The th sparse processing includes: dividing all pixel points in the video frame after the th sparse processing into two categories, and requiring that all neighboring pixel points within the 4-neighborhood of any pixel point are different from the category to which the pixel point belongs; calculating the loss degree when retaining any category; forming the video frame from all pixel points in the category with the smallest loss degree, and taking the smallest loss degree as the loss degree of the video frame ; Calculating the loss degree includes: calculating the contribution coefficient of each neighborhood position within the 4-neighborhood during the th sparse processing according to the gray-scale difference between all pixel points in the non-target category of the two categories and their neighboring pixel points within the 4-neighborhood; obtaining a restored frame of the same size as the video frame through multiple restorations based on the contribution coefficients of each neighborhood position within the 4-neighborhood during the previous th sparse processing and the gray-scale values of all pixel points in the target category, and calculating the loss degree when retaining the target category according to the difference in the gray-scale values of the pixel points in the restored frame and the video frame .

2. The method for optimizing the acquisition of factory videos based on artificial intelligence according to claim 1, wherein, The loss tolerance of the sparse processing , is the monitoring degree corresponding to the monitoring level of the monitoring position, is the monitoring degree corresponding to the monitoring level of the acquisition period of the monitoring video, is the upper limit of the loss tolerance; The monitoring levels include three levels, namely high level, medium level, and low level, and the monitoring degrees corresponding to the three monitoring levels are respectively , , , and .

3. The method for optimizing the acquisition of factory videos based on artificial intelligence according to claim 1, characterized in that, The step of dividing all pixel points in the video frame after the th sparse processing into two categories requires that all neighboring pixel points within the 4-neighborhood of any pixel point are different from the category to which the pixel point belongs, including: For all pixel points in a video frame divide the pixel points where the row and column are odd rows and odd columns and the pixel points where the row and column are even rows and even columns into the first category; divide the pixel points where the row and column are odd rows and even columns and the pixel points where the row and column are even rows and odd columns into the second category.

4. The method for optimizing the acquisition of factory videos based on artificial intelligence according to claim 1, wherein Calculating the contribution coefficient of each neighborhood position in the 4-neighborhood during the th sparse processing, including: The contribution coefficient of the neighborhood position within the 4-neighborhood during the th sparse processing; in the formula, is the average gray-level difference of the neighborhood pixel points at the neighborhood position during the th sparse processing, and , equals , , , sum.

5. The method for optimizing the acquisition of factory videos based on artificial intelligence according to claim 1, wherein, The restored frames obtained through multiple restorations and having the same size as the video frames, including: ​ Set a blank image with the same size as the video frame and take values from 1 to 1; According to the positions of the retained pixel points in the video frame set all the pixel points in the restored frame in the image to obtain the image to be restored ; By the contribution coefficients of each neighborhood position in the 4-neighborhood during the th sparse processing, the gray values of all neighborhood pixels of each blank pixel in the image to be restored are weighted and summed to obtain the gray value of each blank pixel in the image to be restored , and then the gray value of each blank pixel in the image to be restored is obtained, and further the th restored frame is obtained ; Until Stop when it equals 1 and obtain the restored frame , the restored frame is the same size as the video frame .

6. The method for optimizing the acquisition of factory videos based on artificial intelligence according to claim 5, characterized in that The gray values of the blank pixel points in the obtained image to be restored include: ; Wherein, is the gray value of a blank pixel,[ is the number of neighborhood positions where neighborhood pixels exist within the 4-neighborhood of the blank pixel,[ is the th neighborhood position where neighborhood pixels exist within the 4-neighborhood of the blank pixel,[ contribution coefficient of the,[ is the gray value of the th neighborhood pixel existing within the 4-neighborhood of the blank pixel,[ is the th neighborhood position where no neighborhood pixel exists within the 4-neighborhood of the blank pixel,[ contribution coefficient of the,[ is the average value of the gray values of all neighborhood pixels existing within the 4-neighborhood of the blank pixel,[ represents rounding down.[ 7. The method for optimizing the acquisition of factory videos based on artificial intelligence according to claim 5, wherein, The positions of the pixel points retained in the video frame are determined according to the retention class in the video frame: when the retention class is the first class, the positions of the pixel points retained refer to the positions where the row and column are odd rows and odd columns and where the row and column are even rows and even columns; when the retention class is the second class, the positions of the pixel points retained refer to the positions where the row and column are odd rows and even columns and where the row and column are even rows and odd columns.

8. The method for optimizing the acquisition of factory videos based on artificial intelligence according to claim 1, wherein, Based on the difference in the grayscale values of pixel points in the restored frame and the video frame calculating the loss degree when retaining the target class, including: Restore the frame The mean of the differences between the gray values of all pixel points in the and all pixel points in the video frame is used as the loss degree when retaining the target class.

9. The method for optimizing the acquisition of factory videos based on artificial intelligence according to claim 1, characterized in that, The method further includes: The number of times of sparse processing during the process of obtaining the optimized acquisition result The retained classes in the video frames after each sparse processing and the contribution coefficients of each neighborhood position in the 4-neighborhood during each sparse processing are used as supplementary information and transmitted.

10. The method for optimizing the acquisition of factory videos based on artificial intelligence according to claim 9, wherein, The method further includes: at the monitoring center, according to the optimized acquisition result and supplementary information, obtaining video frames for display through multiple restorations, including: setting a blank image with a size equal to the blank image , as the intermediate image of the size, the intermediate image refers to the optimized acquisition result; determining the positions of the retained pixel points in the video frame after the th sparse processing according to the retained classes in the video frame after the th sparse processing, for setting the pixel points in the intermediate image in the blank image ; obtaining the gray values of the remaining blank pixel points by weighted summing the gray values of all neighboring pixel points of the remaining blank pixel points in according to the contribution coefficients of the neighboring positions in the 4-neighborhood during the th sparse processing, obtaining the gray values of the remaining blank pixel points, and obtaining the th intermediate image ; Taking from to 1 until obtaining the intermediate image , as the video frame for display.

Citation Information

Patent Citations

  • Video compression method and system for security and protection monitoring and medium

    CN115294409A

  • Computer network communication data processing method and system based on artificial intelligence

    CN119834933A