AI robot video information extraction system and extraction method
Through preprocessing and optical flow vector calculation combined with convolutional network model, visual neural network distillation training is used to solve the accuracy and real-time problems of moving objects recognition in AI robot video information processing, and improve the recognition accuracy and obstacle avoidance performance.
Patent Information
- Application Number
- CN202510425081.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-22
AI Technical Summary
In the prior art, in the processing of video information of AI robots, there is a contradiction between low recognition accuracy of moving objects, real-time and recognition accuracy, and it is difficult to effectively avoid obstacles in complex environments.
The preprocessing module, optical flow vector calculation module and convolutional network model are used to extract information by combining optical flow vectors and image sequences, and distillation training is used for distillation training to reduce the computational complexity and improve recognition capabilities.
It improves the accuracy and real-time recognition of moving objects by AI robots in complex environments, enhances obstacle avoidance performance, and reduces computing resource consumption.
Smart Images

Figure CN120355752A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image information processing. Specifically, it relates to an AI robot video information extraction system and an extraction method. Background Art
[0002] In the field of AI robots, the recognition of moving objects in video information is one of the key technologies for realizing intelligent obstacle avoidance. In practical applications, not only is it required to accurately recognize moving objects, but also good timeliness is required, that is, it is required to recognize moving objects within a short time.
[0003] The existing technology mainly uses a convolutional neural network to process video information to extract moving objects. This technology reduces the processing dimension of video information by convolving image information, thereby achieving efficient extraction of image information. However, the technical defects of this existing technology are as follows:
[0004] First, the accuracy of directly using the basic convolutional neural network architecture to extract moving objects in video information is relatively low. Although the recognition accuracy can be improved to a certain extent by increasing the number of hidden layers and the number of training iterations of the convolutional neural network, this method is only effective for scenarios within the coverage of training samples.
[0005] Secondly, in practical application scenarios, the working environment of AI robots is highly complex and unpredictable, and the types of objects and motion characteristics contained in video information are extremely rich. Relying solely on annotating the shape characteristics of moving objects in video information for recognition, it is difficult to accurately extract relevant information of moving objects from video information, resulting in insufficient extraction accuracy.
[0006] In addition, there is a contradiction between recognition accuracy and real-time performance in the existing technology: improving recognition accuracy requires a more complex network structure, but it will increase computational latency; while pursuing real-time performance may sacrifice recognition accuracy. This technical contradiction restricts the obstacle avoidance performance of AI robots in dynamic environments. Summary of the Invention
[0007] This section of the content of this application is used to briefly introduce ideas, which will be described in detail in the following detailed implementation section. This section of the content of this application is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0008] As the first aspect of this application, to solve the technical problems mentioned in the above background art section, some embodiments of this application provide an AI robot video information extraction system, including:
[0009] A preprocessing module that preprocesses video information to obtain an image sequence;
[0010] An optical flow vector calculation module that calculates the optical flow vector matrix of adjacent pictures in an image sequence to obtain an optical flow vector sequence corresponding to the image sequence;
[0011] A convolutional network model that inputs the image sequence and the optical flow vector sequence into the convolutional network model to obtain feature information related to moving objects in the image sequence;
[0012] Among them, the convolutional network model is trained by a distillation device;
[0013] The distillation device includes a teacher model; the teacher model is a visual neural network model.
[0014] Compared with the method of using the types of entity information in video information to judge moving objects, in the technical solution provided by this application, the convolutional network model not only inputs the image sequence, but also synchronously inputs the optical flow vector sequence. In this way, the convolutional network model can further use the optical flow vector to capture the moving information in the image sequence, thereby increasing the ability to capture moving objects. And, in order to further increase the ability to capture moving objects, the convolutional network model is also trained by distillation with a visual neural network model. Therefore, even when the structure of the convolutional neural network model is simple, it also has the recognition ability of a large visual neural network model and can accurately extract the feature information of moving objects.
[0015] When extracting moving objects in video information, in order to improve the processing efficiency, it is necessary to reduce the frame rate of the video information to reduce the number of pictures to be processed. However, reducing the picture frame rate may cause the loss of entity information of moving objects in the video information. To solve this technical problem, this application provides the following technical solution:
[0016] Further, the preprocessing module includes:
[0017] A video information acquisition unit for acquiring video information;
[0018] A grouping unit that obtains the sum of the values of any channel of each picture in the video information, and takes the picture sequences with the same sum of values and adjacent to each other as a picture group;
[0019] A frame extraction unit that extracts only one picture for each picture group to generate a picture sequence;
[0020] A preprocessing unit that preprocesses each picture in the picture sequence.
[0021] In the technical solution provided by this application, when extracting picture information, instead of directly extracting frames from the picture sequence, pictures with the same sum of values in any channel are grouped as a picture group, and frames are extracted from this picture group. Since the sum of values in any channel of such pictures is equal, the content in these pictures basically does not change, so using them for frame extraction will not cause the loss of the entity information of moving objects. In this way, the number of pictures to be processed is reduced.
[0022] Further, the preprocessing includes filtering, denoising, and normalization processing.
[0023] Filtering, denoising, and normalization processing of the picture information can process each picture in the video information into the same format, thereby increasing the system's processing ability for pictures.
[0024] When calculating the optical flow vector, due to the extremely large amount of calculation, it may be impossible to quickly extract the information of moving objects in the video information. For this reason, this application provides the following technical solution:
[0025] Further, the optical flow vector calculation module includes:
[0026] A picture grouping unit that divides adjacent pictures in the picture sequence into several optical flow calculation groups;
[0027] A GPU acceleration unit, which is provided with multiple GPU calculation chips for parallelly calculating the optical flow vector matrix in each optical flow calculation group;
[0028] An optical flow vector combination unit that combines all the optical flow vector matrices into an optical flow vector sequence.
[0029] In the technical solution provided by this application, multiple GPU calculation chips are used to parallelly calculate the optical flow vector matrix of each optical flow calculation group, and the ability of the GPU calculation chip to handle simple calculations is utilized to quickly complete the calculation of the optical flow vector.
[0030] When calculating the optical flow vector, it is necessary to traverse each pixel point, so the amount of calculation is large and a large amount of computing resources are consumed. For this problem, this application provides the following technical solution:
[0031] Further, the GPU acceleration unit divides the picture into several regions and calculates the average optical flow vector of each region P; where P is the index of the region.
[0032] In the technical solution provided by this application, calculating the average optical flow vector of each region can reduce the consumption of computing resources and can also reflect the optical flow information of the picture.
[0033] When dealing with small target objects, in order to increase the recognition accuracy, it is necessary to increase the accuracy of the optical flow vector as much as possible. However, too precise calculation of the optical flow vector will increase the calculation amount of the optical flow vector and affect the information processing efficiency. Therefore, the present application provides the following technical solutions:
[0034] The GPU acceleration unit calculates the average optical flow vector of each region based on the following steps:
[0035] S1: Reduce the resolution of region P by 1 / 4 to obtain the L2 sampling layer;
[0036] S2: Calculate the optical flow vector of the L2 sampling layer. If the optical flow vector of the L2 sampling layer is less than the second preset value, take the optical flow vector of the L2 sampling layer as the sampling layer of region P. If it is greater than the second preset value, perform the subsequent steps:
[0037] S3: Reduce the resolution of region P by 1 / 2 to obtain the L1 sampling layer;
[0038] S4: Calculate the optical flow vector of the L1 sampling layer. If the optical flow vector of the L1 sampling layer is less than the first preset value, take the optical flow vector of the L1 sampling layer as the sampling layer of region P. If it is greater than the first preset value, perform the subsequent steps:
[0039] S5: Calculate the optical flow vector of region P.
[0040] In the technical solutions provided by the present application, when calculating the optical flow vector of each region, a hierarchical calculation method is adopted. First, calculate the optical flow vector under the condition of low resolution. If the optical flow vector is greater than the preset value, it indicates that there is abnormal information. Therefore, to further improve the calculation efficiency of the optical flow vector, this solution calculates the optical flow vectors of region P at different resolutions in a hierarchical manner, and then can obtain an accurate optical flow vector matrix with less computing resources.
[0041] When the optical flow vector and video information are input into the convolutional neural network model twice, it will cause the convolutional neural network model to be unable to understand the relationship between two adjacent input information, and thus unable to extract moving objects in the video information by combining the optical flow vector. Therefore, the present application provides the following technical solutions:
[0042] Further, the AI robot video information extraction system further includes an information integration module;
[0043] The information integration module converts each picture in the picture sequence into a gray value, and takes the gray value as the first-dimensional information;
[0044] For the optical flow vector matrix corresponding to each picture, take the optical flow component in the horizontal direction in the optical flow vector matrix as the second-dimensional information;
[0045] For the optical flow vector matrix corresponding to each picture, the optical flow components in the longitudinal direction in the optical flow vector matrix are used as the third-dimensional information;
[0046] The first-dimensional information, the second-dimensional information, and the third-dimensional information are combined into an information matrix T, and the information matrix T is input into the convolutional network model.
[0047] In the technical solution provided by this application, the optical flow vector and the gray value in the picture are fused together as an information matrix. Therefore, the optical flow vector and the video information are actually input into the convolutional network model synchronously. Therefore, when the convolutional network model extracts the information in the information matrix T, it can effectively identify the feature information of the moving object therein, and then accurately extract the moving object and the moving trajectory of the moving object.
[0048] When existing convolutional neural networks process data information, they generally process single-modal data, that is, the underlying logic of the convolutional neural network is an AND-OR gate. This results in that when the convolutional neural network processes data, if the number of network layers of the convolutional layer is not increased, the fineness of information processing is not high, and the implicit features related to motion information cannot be accurately extracted. For this reason, this application provides the following technical solutions:
[0049] Further, the convolutional network model includes:
[0050] An input layer that converts the information matrix into multi-modal data;
[0051] A convolutional layer for extracting implicit features in the multi-modal data;
[0052] A pooling layer for reducing the data dimension;
[0053] An output layer for outputting motion information related to the moving object, where the motion information includes the range of the moving object and the moving direction of the moving object.
[0054] In this solution, converting the information matrix into multi-modal data can enrich the types of information in practice, and then the convolutional layer can better extract the implicit features related to motion information from the multi-modal data.
[0055] Further, the visual neural network model includes:
[0056] A brightness information extraction layer: inputting video information and extracting the brightness change information of pixel points from the video information;
[0057] An excitation layer: extracting visual excitation information from the brightness change information;
[0058] An inhibition layer: extracting visual inhibition information from the brightness change information;
[0059] Information integration layer: Receives visual excitation information and visual inhibition information, and generates fused features;
[0060] Information output layer: Receives the fused features and generates motion information of the moving object based on the fused features.
[0061] In the technical solution provided by this application, a more complex visual neural network model is used as the teacher model. The teacher model can analyze visual excitation information and visual inhibition information based on the brightness change information, and then accurately extract the motion information of the moving object. In this way, using a high-precision network model to train the convolutional network model can reduce the complexity of the convolutional network model, reduce the computational difficulty of the convolutional network model, and at the same time increase the prediction accuracy of the convolutional network model.
[0062] When calculating the optical flow vector, if the global optical flow vector is directly calculated based on the brightness, it will cause the finally calculated optical flow vector to tend to be locally optimal, resulting in the globally weak optical flow vector generated finally and unable to represent the real optical flow state. For this reason, this application provides the following technical solution:
[0063] Further, the calculation method of the optical flow vector is as follows:
[0064] Step1: Convert the two pictures for which the optical flow vector needs to be calculated into grayscale images I1 and I2, and calculate the pixel difference I t (x, y);
[0065] Step2: Calculate the gradients of the grayscale images I1 and I2 in the x and y directions;
[0066] The calculation formula of the gradient is:
[0067]
[0068] I t (x, y) = I2(x, y, t + 1) - I1(x, y, t);
[0069] Among them, I x Indicates the spatial gradient of the image in the x direction, I y Represents the spatial gradient of the image in the y direction, I t Represents the temporal gradient of the grayscale images I1 and I2; K x Represents the Sobel kernel in the x direction, K y Represents the Sobel kernel in the y direction, t represents time, and i and j respectively represent the row and column offsets of the Sobel kernel.
[0070] Step3: Set the optical flow field to 0, and set the iterative parameters smoothing weight λ, maximum iteration number R, and convergence threshold σ;
[0071] u o (x, y) = 0, v o (x, y) = 0;
[0072] u o (x, y) represents the initial optical flow component in the x - direction of grayscale image I1 and grayscale image I2, v o (x, y) represents the initial optical flow component in the y - direction of grayscale image I1 and grayscale image I2;
[0073] Step3: Iterate the optical flow field to obtain the optical flow component of each pixel point;
[0074] The iteration process is as follows:
[0075] For each pixel (x, y), calculate the mean of its neighborhood;
[0076]
[0077] and respectively represent the average optical flow in the x - direction and y - direction of the pixel point (x, y) in the neighborhood;
[0078] Calculate the brightness residual W;
[0079]
[0080] Update the optical flow component;
[0081]
[0082] where r represents the number of iterations, u r+1 (x, y) and v r+1 (x, y) respectively represent the optical flow components in the x - direction and y - direction of the pixel point (x, y) at the (r + 1)-th iteration.
[0083] When the maximum number of iterations is reached or the convergence threshold is reached, stop updating the optical flow vector.
[0084] In the technical solution provided by this application, when calculating the optical flow vector, it is not directly calculated once, but multiple iterations of optical flow are performed. Therefore, the finally generated optical flow vector information conforms to the actual optical flow field, and the accuracy of the optical flow information is higher.
[0085] As the second aspect of this application, this application provides an AI robot video information extraction method, which uses the described AI robot video information extraction system to extract the motion information in the video. Description of the Drawings
[0086] The accompanying drawings, which form a part of this application, are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more apparent. The schematic embodiments and their descriptions in the accompanying drawings are used to explain this application and do not constitute an improper limitation to this application.
[0087] In addition, throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the elements and components are not necessarily drawn to scale.
[0088] In the drawings:
[0089] Figure 1 It is a schematic structural diagram of an AI robot video information extraction system.
[0090] Figure 2 It is a schematic structural diagram of a preprocessing module.
[0091] Figure 3 It is a schematic structural diagram of an optical flow vector calculation module.
[0092] Figure 4 It is a schematic structural diagram of a convolutional layer. Detailed implementation manners
[0093] The embodiments of this application will be described in more detail below with reference to the accompanying drawings. Although some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand this application. It should be understood that the drawings and embodiments of this application are only for exemplary purposes and are not used to limit the protection scope of this application.
[0094] In addition, it should be noted that for the convenience of description, only the parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in this application and the features in the embodiments can be combined with each other.
[0095] This application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0096] The convolutional neural network is an efficient model for extracting image information. It can accurately extract the implicit information in images and is suitable for extracting entity information in images. For example, it can identify various static objects such as vehicles and people in images. When applying the convolutional neural network model to video information extraction, a preliminary inference of the motion state is made based on the identified object type. For example, when a specific category such as a vehicle is recognized, it can be inferred that it may be a moving object. That is to say, the convolutional neural network model is mainly suitable for the category recognition of static objects and is difficult to accurately identify and extract the information of objects in a moving state. Therefore, it is difficult for the convolutional neural network model to extract the feature information of moving objects in video information.
[0097] The recognition ability of the convolutional network model depends to a large extent on the training situation. In this solution, the optical flow vector is integrated into the image information, and the optical flow vector and the image information are input into the convolutional network model together. The convolutional network model can utilize the implicit information of the optical flow vector and its ability to recognize entity information to further understand the moving objects in the video information, so as to accurately extract the features of the moving objects in the video information. At the same time, in order to accurately train the convolutional network model, this solution also uses a large visual neural network model as the teacher model to perform distillation training on the convolutional network model, assisting the convolutional network model to integrate the connection between the image information and the optical flow vector information during training, so as to replace the large visual neural network model with a convolutional network model with a simple structure and low construction cost. The specific solution is as follows:
[0098] Example 1: Refer to Figures 1 to 3 , the AI robot video information extraction system includes:
[0099] A preprocessing module for obtaining video information, preprocessing the video information to obtain an image sequence;
[0100] An optical flow vector calculation module for calculating the optical flow vector matrix of adjacent pictures in the image sequence to obtain an optical flow vector sequence corresponding to the image sequence;
[0101] A convolutional network model, the image sequence and the optical flow vector sequence are input into the convolutional network model to obtain the feature information related to the moving objects in the image sequence;
[0102] The convolutional network model is trained through a distillation device. The distillation device includes a teacher model. The teacher model is a visual neural network model.
[0103] The feature information in this solution is the moving objects in the video information and the overall offset direction of the moving objects.
[0104] The key of this application lies in using a convolutional network model to extract the motion information from video information. When dealing with picture information, the convolutional network model has good entity information extraction ability and can accurately extract each entity information in the picture information. In order to extract the motion directions of each entity in the picture information, an optical flow vector is further added. The optical flow vector is also input into the convolutional network model, so that the convolutional network model can further obtain the motion directions of each entity information on the basis of extracting the entity information, that is, obtain the motion information of the moving objects in the video information.
[0105] In order to enable the convolutional network model to organically integrate the picture information and the optical flow vector information, this application uses a teacher model to perform distillation training on the convolutional network model, so as to accurately extract the motion information of the moving objects.
[0106] The teacher model can be a visual neural network model. However, the visual neural network model requires a large amount of computing resources and is not conducive to being deployed in small terminal devices. Therefore, in this solution, a convolutional network model is used as the student model of the visual neural network model. Taking advantage of the characteristics that the convolutional network model can be miniaturized and deployed, a miniaturized convolutional network model is used as the motion information extraction model, and then the visual neural network model is used to train the convolutional network model. In this way, both the extraction efficiency of the motion information can be guaranteed and the consumption of computing resources can be reduced.
[0107] When extracting the moving objects in the video information, the higher the frame rate of the video information, the higher the accuracy of the motion information extraction. However, too high a frame rate will lead to an increase in the number of pictures to be processed. Therefore, this application provides the following technical solutions:
[0108] Furthermore, the preprocessing module includes: a video information acquisition unit, a grouping unit, a frame extraction unit, and a preprocessing unit. Among them:
[0109] The video information acquisition unit is used to acquire video information;
[0110] The grouping unit obtains the sum of the values of any channel of each picture in the video information, and takes the adjacent picture sequences with the same sum of values as a picture group;
[0111] The frame extraction unit only extracts one picture from each picture group to generate a picture sequence.
[0112] This solution selectively extracts frames from video information. Generally, the pictures in video information are RGB format pictures. If in a certain video, there are no moving objects in a continuous segment of video (multiple consecutive frames), then the pictures in the video information will not change. Thus, the sum of pixel values in any color channel will also not change. Based on this characteristic, consecutive frames with the same content are grouped into a picture group, and one picture is extracted from the picture group, while the remaining redundant pictures are deleted. In this way, the video information will be selectively frame-extracted, and a large number of invalid pictures with the same information content are deleted, only the valid picture information containing the information of moving objects in the picture sequence is retained. At the same time, this screening method only needs to add up all the pixel values of a certain channel, without the need to conduct detailed comparison for each picture. Therefore, the computational complexity is not high, and it can determine whether pictures are the same with relatively little computing resources. In this solution, the selection of any channel is related to the total number in the picture. If the value of the R channel in the picture is the highest, then the R channel is selected as the basis for dividing the picture group; if the value of the G channel is the highest, then the G channel is selected as the basis for dividing the picture group.
[0113] The preprocessing unit preprocesses each picture in the picture sequence. The preprocessing mainly includes filtering, denoising, and normalization of the pictures. The preprocessing unit belongs to the prior art, and there is a lot of teaching in the prior art, which will not be elaborated here.
[0114] The optical flow vector is actually the moving direction of pixel points between two adjacent pictures. If the entity information in two pictures does not change, then the optical flow vector corresponding to each pixel point is 0. On the contrary, if there is certain moving information in the entity information, the optical flow vector will describe the moving information of each pixel point. In this way, the convolutional neural network only needs to segment the entities in the picture and then fuse the segmented entities with the optical flow vector to analyze the moving entity information in the video information.
[0115] In this solution, an optical flow vector matrix is calculated for two adjacent pictures, and then the optical flow vector matrix is corresponding to the latter picture of the two pictures. In this way, all optical flow vector matrices are calculated. The calculated optical flow vector matrices are arranged in sequence to obtain an optical flow vector sequence. After the optical flow vector sequence is aligned with the picture sequence, the picture information can be corresponding to the optical flow vector matrix.
[0116] When calculating the optical flow vector, a large amount of resources are consumed. In order to reduce the computing resources, this application adopts the following method:
[0117] The optical flow vector calculation module includes: an image grouping unit, a GPU acceleration unit, and an optical flow vector combination unit. Among them, the image grouping unit divides adjacent images in the image sequence into several optical flow calculation groups. The image grouping unit groups the collected image sequence, and each optical flow calculation group calculates an optical flow vector matrix.
[0118] The GPU acceleration unit sets multiple GPU calculation chips, and these GPU calculation chips are set in parallel. Therefore, the calculation of each optical flow vector matrix is carried out in parallel, which can increase the generation speed of the optical flow vector matrix.
[0119] The optical flow vector combination unit combines all the optical flow vector matrices into an optical flow vector sequence. In this way, after the GPU calculation chips calculate the optical flow vector matrices, arranging these optical flow addition matrices in order can obtain the optical flow vector sequence.
[0120] When calculating the optical vector matrix, in order to improve the calculation efficiency, the present application provides the following two solutions:
[0121] Solution 1:
[0122] The GPU acceleration unit divides the image into several regions and calculates the average optical flow vector of each region. For example, the image is equally divided into 9 regions, and these 9 regions are regarded as a pixel point to calculate the overall optical flow vector of the region, so as to reduce the calculation amount of the optical flow vector. That is, first divide the image into several regions, for example, divide it into 9 regions, calculate the average optical flow vector of each region, and then combine the average optical flow vectors of these regions to obtain the optical flow vector matrix.
[0123] However, the calculation accuracy of this optical flow vector is not high. For this reason, the present application further provides the following solution:
[0124] Solution 2: The GPU acceleration unit divides the image into several regions, and the GPU acceleration unit calculates the average optical flow vector of each region based on the following steps. Among them, the Pth region is region P, and P is the index of the region.
[0125] Specifically:
[0126] S1: Reduce the resolution of region P by 1 / 4 to obtain the L2 sampling layer.
[0127] S1 is to reduce the resolution of region P by 1 / 4. For example, if the resolution of the region is 400*400, after reducing by 1 / 4, the resolution becomes 100*100. How to reduce the resolution is the prior art and will not be elaborated here. In practice, the average pixel value of 4 pixel grids is used as the pixel grid after reducing the resolution, so that the effect of reducing the resolution can be achieved.
[0128] S2: Calculate the optical flow vector of the L2 sampling layer. If the optical flow vector of the L2 sampling layer is less than the second preset value, use the optical flow vector of the L2 sampling layer as the sampling layer of region P; if it is greater than the second preset value, perform the subsequent steps:
[0129] After obtaining the L2 sampling layer, the resolution of the region decreases, and the computational amount of the corresponding optical flow vector also decreases. If the total value of the optical flow vectors within the entire region is less than the second preset value, it indicates that the change in the optical flow vectors of this region is not significant. Therefore, there is no need to further increase the computational accuracy of the optical flow vectors; otherwise, it is necessary to further increase the computational accuracy of the optical flow vectors and then perform the subsequent steps.
[0130] To calculate the total number of optical flow vectors, only need to add the absolute values of the movements in the x and y directions of the optical flow vectors, and then compare with the second preset value.
[0131] S3: Reduce the resolution of region P by 1 / 2 to obtain the L1 sampling layer;
[0132] S4: Calculate the optical flow vector of the L1 sampling layer. If the optical flow vector of the L1 sampling layer is less than the first preset value, use the optical flow vector of the L2 sampling layer as the sampling layer of region P; if it is greater than the first preset value, perform the subsequent steps:
[0133] The logic of S3 and S4 is the same as the foregoing and will not be elaborated here.
[0134] S5: Calculate the optical flow vector of region P.
[0135] If the optical flow vectors of both the L1 sampling layer and the L2 sampling layer are greater than the preset value, it indicates that the optical flow of this region changes violently. Therefore, region P cannot be downsampled.
[0136] In this way, the GPU acceleration unit can quickly and accurately calculate the optical flow vector matrix using this scheme. Therefore, for region P, if there are no moving objects in this region P, when calculating the average optical flow vector of this region, after reducing the clarity of region P by 1 / 4, use the calculated optical flow vector as the average optical flow vector of this region. If there is certain motion information, after reducing the clarity of region P by 1 / 2, use the calculated optical flow vector as the average optical flow vector of this region; if there is a large amount of motion information, it is necessary to strictly calculate the optical flow vector of region P and use this vector as the average optical flow vector.
[0137] The above is the partition calculation logic of the optical flow vectors. When calculating the specific optical flow vectors, after corresponding two pictures to each other, perform the following scheme.
[0138] The calculation method of the optical flow vectors is as follows:
[0139] Step1: Convert the two images for which the optical flow vectors need to be calculated into grayscale images I1 and I2, and calculate the pixel difference I t (x, y).
[0140] In this solution, the grayscale images I1 and I2 can be understood as two regions for which the optical flow vectors need to be calculated.
[0141] Step2: Calculate the gradients of the grayscale images I1 and I2 in the x and y directions;
[0142] The calculation formula for the gradient is:
[0143]
[0144] I t (x, y) = I2(x, y, t + 1) - I1(x, y, t);
[0145] where, I x represents the spatial gradient of the image in the x direction, I y represents the spatial gradient of the image in the y direction, I t represents the temporal gradient of the grayscale images I1 and I2; K x represents the Sobel kernel in the x direction, K y represents the Sobel kernel in the y direction;
[0146] Step3: Set the optical flow field to 0, and set the iterative parameters: smoothing weight λ, maximum number of iterations R, and convergence threshold σ;
[0147] u o (x, y) = 0, v o (x, y) = 0;
[0148] u o (x, y) represents the initial optical flow component of the grayscale images I1 and I2 in the x direction, v o (x, y) represents the initial optical flow component of the grayscale images I1 and I2 in the y direction;
[0149] Step3: Iterate the optical flow field to obtain the optical flow components of each pixel point;
[0150] The iteration process is as follows:
[0151] For each pixel (x, y), calculate the mean value of its neighborhood;
[0152]
[0153] and respectively represent the average optical flow of the pixel point (x, y) in the x direction and y direction neighborhood.
[0154] Calculate the luminance residual W;
[0155]
[0156] Update the optical flow components;
[0157]
[0158] where r represents the number of iterations, and u r+1 (x, y) and v r+1 (x, y) respectively represent the optical flow components in the x - direction and y - direction of the pixel point (x, y) at the (r + 1)-th iteration.
[0159] When the maximum number of iterations is reached, or when the convergence threshold is reached, stop updating the optical flow vector.
[0160] In the technical solution provided by this application, when calculating the optical flow vector, it is not directly calculated in one go, but rather through multiple iterative calculations of the optical flow. The finally generated optical flow vector information conforms to the actual optical flow field, and the accuracy of the optical flow information is higher.
[0161] Using the above - mentioned solution can calculate the optical flow vector of each region. Then, by combining these regions, an optical flow vector matrix can be obtained. However, if the optical flow vector matrix and the picture are alternately input into the convolutional network model, it may cause confusion in the convolutional network model's understanding of the relationship between the optical flow vector matrix and the picture. Therefore, in this application, the optical flow vector matrix and the picture are first fused to form a closely related piece of information (i.e., as an input information). This processing method reduces the difficulty for the convolutional network model to understand these two types of information.
[0162] The AI robot video information extraction system further includes an information integration module. The information integration module converts each picture in the picture sequence into a grayscale value, and uses the grayscale value as the first - dimension information. For the optical flow vector matrix corresponding to each picture, the optical flow components in the horizontal direction of the optical flow vector matrix are used as the second - dimension information. For the optical flow vector matrix corresponding to each picture, the optical flow components in the vertical direction of the optical flow vector matrix are used as the third - dimension information. The first - dimension information, the second - dimension information, and the third - dimension information are combined into an information matrix T, and the information matrix T is input into the convolutional network model.
[0163] T(x, y)=(q1, q2, q3), where T(x, y) represents the information at the position (x, y) in the information matrix T, and q1, q2, q3 respectively represent the first - dimension information, the second - dimension information, and the third - dimension information.
[0164] Therefore, the information matrix T contains the gray value of each pixel point, as well as the optical flow trends (optical flow vectors) of each pixel point in the horizontal and vertical directions.
[0165] The convolutional network model includes: an input layer, a convolutional layer, a pooling layer, and an output layer.
[0166] Among them, the input layer converts the information matrix into multi-modal data ψ.
[0167] The information matrix is three-dimensional information, and the input layer is converted into multi-modal data in the following way:
[0168] Ψ = L(θ1)|0> ⊕ L(θ2)|0> ⊕ L(θ3)|0>;
[0169] θ1 = π(x1 / max[x1]);
[0170] θ2 = π(x2 / max[x2]);
[0171] θ3 = π(x3 / max[x3]);
[0172] Among them, max[x1] represents the maximum value of the information in the first dimension; max[x2] represents the maximum value of the information in the second dimension; max[x3] represents the maximum value of the information in the third dimension; |0> is the Dirac symbol, ⊕ is the tensor product, and L represents the single-bit rotation gate operation;
[0173] The convolutional layer is used to extract the implicit features ψ` in the multi-modal data;
[0174] As Figure 4 shown, the convolutional layer includes a CRz gate and a CRx gate, which are alternately composed. The CRz gate is used to adjust the phase of the multi-modal input, and the CRx gate is used to perform a rotation operation on the adjacent input data.
[0175] Among them, the calculation of the convolution kernel in the convolutional layer is G(θ);
[0176] G(θ) = (CRz(θ1)CRx(θ2))(CRz(θ3)CRx(θ4))(CRz(θ5)
[0177] CRx(θ6));
[0178] ψ` = G(θ)ψ;
[0179] In this way, the convolution kernel in the convolutional layer provided by this application can adjust 6 conditional parameters at the same time, and thus has better feature extraction ability compared with the ordinary convolutional layer.
[0180] The pooling layer is used to reduce the data dimension;
[0181] An output layer for outputting motion information related to a moving object, where the motion information includes the range of the moving object and the moving direction of the moving object.
[0182] In the technical solution provided by this application, the pooling layer and the output layer are the pooling layer and the output layer in an existing convolutional neural network model. For the convolutional layer, this application uses an alternating composition of CRz gates and CRx gates to replace the ordinary convolutional kernel, so that six conditional parameters can be adjusted, thereby facilitating the understanding of higher-dimensional features.
[0183] The visual neural network model includes:
[0184] A brightness information extraction layer: input video information and extract the brightness change information of pixel points from the video information;
[0185] An excitation layer: extract visual excitation information from the brightness change information;
[0186] An inhibition layer: extract visual inhibition information from the brightness change information;
[0187] An information integration layer: receive visual excitation information and visual inhibition information and generate a fused feature;
[0188] An information output layer: receive the fused feature and generate motion information of the moving object based on the fused feature.
[0189] The visual neural network model is a common large-scale information extraction model in the art. For example: the locust visual neural network model (the locust visual neural network model is a bionic computing model inspired by the locust visual system, mainly simulating its efficient dynamic target detection and collision avoidance capabilities).
[0190] The distillation device includes a teacher model. The teacher model is a visual neural network model. The convolutional network model is trained through the distillation device.
[0191] The training process is as follows:
[0192] Step 1: Input the video information into the visual neural network, and then input the information matrix corresponding to the video information into the convolutional network model.
[0193] The result generated by the visual neural network model is used as a soft label, and the convolutional network model outputs a hard label;
[0194] Step 2: Calculate the cross-entropy loss using the hard label, calculate the distillation loss using the soft label, and sum the distillation loss and the cross-entropy loss with weights to obtain the total loss;
[0195] Step 3: Update the parameters of the convolutional network model through backpropagation to minimize the total loss.
[0196] The total loss is the cross-entropy loss function. That is, in this solution, when training the convolutional network model, the cross-entropy loss function will introduce a distillation loss with the visual neural network model. Therefore, when training the convolutional model, it can be affected by the visual neural network model to further find the correlation between the optical flow vector and the gray value.
[0197] Embodiment 2: The present application provides an AI robot video information extraction method, and uses the described AI robot video information extraction system to extract the motion information in the video.
[0198] The above description is only some preferred embodiments of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the embodiments of the present application.
Claims
1. An AI robot video information extraction system, characterized in that, Including: A preprocessing module that preprocesses video information to obtain an image sequence; An optical flow vector calculation module that calculates the optical flow vector matrix of adjacent pictures in the image sequence to obtain an optical flow vector sequence corresponding to the image sequence; A convolutional network model that inputs the image sequence and the optical flow vector sequence into the convolutional network model to obtain feature information related to moving objects in the image sequence; Among them, the convolutional network model is trained by a distillation device; The distillation device includes a teacher model; the teacher model is a visual neural network model.
2. The AI robot video information extraction system according to claim 1, characterized in that, The preprocessing module includes: A video information acquisition unit for acquiring video information; A grouping unit that obtains the sum of the values of any channel of each picture in the video information, and takes the picture sequences with the same sum of values and adjacent to each other as a picture group; A frame extraction unit that extracts only one picture for each picture group to generate an image sequence; A preprocessing unit that preprocesses each picture in the image sequence.
3. The AI robot video information extraction system according to claim 1, characterized in that: The preprocessing includes filtering, denoising, and normalization processing.
4. The AI robot video information extraction system according to claim 1, wherein: The optical flow vector calculation module includes: A picture grouping unit that divides adjacent pictures in the image sequence into several optical flow calculation groups; A GPU acceleration unit provided with multiple GPU calculation chips for parallelly calculating the optical flow vector matrix in each optical flow calculation group; An optical flow vector combination unit that combines all the optical flow vector matrices into an optical flow vector sequence.
5. The AI robot video information extraction system according to claim 4, characterized in that: The GPU acceleration unit divides the picture into several regions and calculates the average optical flow vector of each region P; where P is the index of the region.
6. The AI robot video information extraction system according to claim 5, characterized in that The GPU acceleration unit calculates the average optical flow vector of each region based on the following steps: S1: Reduce the resolution of region P by 1 / 4 to obtain an L2 sampling layer; S2: Calculate the optical flow vector of the L2 sampling layer. If the optical flow vector of the L2 sampling layer is less than the second preset value, take the optical flow vector of the L2 sampling layer as the sampling layer of region P. If it is greater than the second preset value, perform the subsequent steps: S3: Reduce the resolution of region P by 1 / 2 to obtain an L1 sampling layer; S4: Calculate the optical flow vector of the L1 sampling layer. If the optical flow vector of the L1 sampling layer is less than the first preset value, take the optical flow vector of the L1 sampling layer as the sampling layer of region P. If it is greater than the first preset value, perform the subsequent steps: S5: Calculate the optical flow vector of region P.
7. The AI robot video information extraction system according to claim 1, wherein, The AI robot video information extraction system further includes an information integration module; The information integration module converts each picture in the image sequence into a gray value and takes the gray value as the first dimension information; For the optical flow vector matrix corresponding to each picture, take the optical flow component in the horizontal direction in the optical flow vector matrix as the second dimension information; For the optical flow vector matrix corresponding to each picture, take the optical flow component in the vertical direction in the optical flow vector matrix as the third dimension information; Combine the first dimension information, the second dimension information, and the third dimension information into an information matrix T and input the information matrix T into the convolutional network model.
8. The AI robot video information extraction system according to claim 7, wherein, The convolutional network model includes: An input layer that converts the information matrix into multimodal data; A convolutional layer for extracting implicit features in the multimodal data; A pooling layer for reducing the data dimension; An output layer for outputting motion information related to a moving object, where the motion information includes the range of the moving object and the moving direction of the moving object.
9. The AI robot video information extraction system according to claim 8, wherein, The visual neural network model includes: A brightness information extraction layer: inputting video information and extracting the brightness change information of pixel points from the video information; An excitation layer: extracting visual excitation information from the brightness change information; An inhibition layer: extracting visual inhibition information from the brightness change information; An information integration layer: receiving the visual excitation information and the visual inhibition information and generating a fusion feature; An information output layer: receiving the fusion feature and generating the motion information of the moving object based on the fusion feature.
10. An AI robot video information extraction method, characterized in that: The motion object information in the video is extracted by using the AI robot video information extraction system described in any one of claims 1 to 9.