Visual data processing method and device and storage medium

By separating the visual data from high-frequency and low-frequency features and processing it using the channel dimension attention module, a visual image code stream suitable for different task types is solved, and the problem that the existing technology cannot be applied to various task scenarios is improved, and the reliability of task execution is improved.

CN119942392APending Publication Date: 2025-05-06ZTE CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311409639.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-27
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing machine vision encoding schemes cannot be applied to various task scenarios, affecting the reliability of task execution.

Method used

The visual data is processed through an encoder, high-frequency characteristic data and low-frequency characteristic data are separated, and the channel dimension attention module is used to process it, and the task target data is determined according to the task type for encoding processing, and a visual image code stream is generated.

Benefits of technology

The task-based visual data processing is realized, the reliability of task execution is improved, and it can adapt to the needs of different task scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942392A_ABST
    Figure CN119942392A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a visual data processing method and device and a storage medium, and belongs to the technical field of data processing. The method comprises the steps that acquired visual data are processed through an encoder, high-frequency feature data and low-frequency feature data corresponding to the visual data are obtained, and the encoder comprises an octave convolution module; inputting the high-frequency feature data into a channel dimension attention module for processing to obtain first high-frequency feature component data and second high-frequency feature component data; determining task target data based on the low-frequency feature data, the first high-frequency feature component data and the second high-frequency feature component data according to the task type of the current visual image task, and encoding the task target data to obtain a visual image code stream; and transmitting the visual image code stream to a decoding end for decoding to obtain decoded data, and completing the current visual image task according to the decoded data. According to the technical scheme, the reliability of task execution is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a visual data processing method, device and storage medium. Background Art

[0002] With the rapid development of computer vision and machine learning, the application scope of machine vision coding is continuously expanding. Smart cities, smart homes, industrial automation, autonomous driving and other fields all require machines to process image and video data to complete related computer vision tasks. Unlike traditional video coding, the field of machine vision coding is committed to developing coding schemes for both human and machine vision. Currently, individual machine vision tasks can be completed through feature coding. However, for scenarios requiring human observation or evidence collection, image reconstruction is also required. However, image reconstruction tasks and machine vision tasks have different quality requirements and priorities. Therefore, machine vision coding is not applicable to all task scenarios, thus affecting the reliability of task execution.

[0003] Therefore, how to realize task-based visual data processing to improve the reliability of task execution has become an urgent problem to be solved. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to provide a visual data processing method, device and storage medium, aiming to realize task-based visual data processing to improve the reliability of task execution.

[0005] In a first aspect, an embodiment of the present application provides a method for processing visual data, the method comprising:

[0006] Processing the acquired visual data through an encoder to obtain high-frequency feature data and low-frequency feature data corresponding to the visual data, wherein the encoder includes an octave convolution module;

[0007] Inputting the high-frequency feature data into the channel dimension attention module for processing to obtain first high-frequency feature component data and second high-frequency feature component data;

[0008] According to the task type of the current visual image task, determining task target data based on the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data, encoding the task target data to obtain a visual image code stream;

[0009] The visual image code stream is transmitted to a decoding end for decoding to obtain decoded data, and the current visual image task is completed according to the decoded data.

[0010] In a second aspect, an embodiment of the present application provides a visual data processing device, the visual data processing device comprising:

[0011] A processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing connection and communication between the processor and the memory, wherein when the computer program is executed by the processor, the steps of the visual data processing method as described above are realized.

[0012] In a third aspect, an embodiment of the present application provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the visual data processing method as described above.

[0013] The embodiments of the present application provide a visual data processing method, device and storage medium, which processes the acquired visual data through an encoder to obtain high-frequency feature data and low-frequency feature data corresponding to the visual data, inputs the high-frequency feature data into the channel dimension attention module for processing, obtains first high-frequency feature component data and second high-frequency feature component data, and determines task target data based on the low-frequency feature data, the first high-frequency feature component data and the second high-frequency feature component data according to the task type of the current visual image task, encodes the task target data to obtain a visual image code stream, transmits the visual image code stream to a decoding end for decoding, obtains decoded data, and completes the current visual image task according to the decoded data. That is, for various tasks, the tasks are completed by performing different data processing methods on the visual data, thereby improving the reliability of task execution. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0015] Figure 1 A flowchart of a visual data processing method provided by an embodiment of the present invention;

[0016] Figure 2 A schematic diagram of the structure of an encoder provided by an embodiment of the present invention;

[0017] Figure 3 This is a schematic diagram of the octave convolution module structure;

[0018] Figure 4 A schematic diagram of a process for obtaining high-frequency feature data and low-frequency feature data corresponding to the visual data through an encoder processing provided by an embodiment of the present invention;

[0019] Figure 5 A schematic diagram of the self-calibration convolutional network module structure;

[0020] Figure 6 This is a schematic diagram of the channel dimension attention module structure;

[0021] Figure 7 A schematic diagram of a process for obtaining first high-frequency characteristic component data and second high-frequency characteristic component data provided by an embodiment of the present invention;

[0022] Figure 8 A schematic diagram of obtaining first high-frequency feature component data and second high-frequency feature component data based on a channel-dimensional attention module provided by an embodiment of the present invention;

[0023] Figure 9 A schematic diagram of the structure of a hyperparameter encoder provided by an embodiment of the present invention;

[0024] Figure 10 A schematic diagram of the structure of an encoding terminal provided by an embodiment of the present invention;

[0025] Figure 11 Schematic diagram of the network structure for multi-scale feature reconstruction;

[0026] Figure 12 It is the schematic diagram of the structure of FT1 and FT2 modules;

[0027] Figure 13 A schematic diagram of a decoding end structure provided by an embodiment of the present invention;

[0028] Figure 14 A schematic diagram of a network overall framework structure provided by an embodiment of the present invention;

[0029] Figure 15 A schematic diagram of a process flow for executing various tasks provided in an embodiment of the present invention;

[0030] Figure 16 A schematic block diagram of the structure of a visual data processing device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0031] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0032] The flowcharts shown in the accompanying drawings are for illustrative purposes only and do not necessarily include all contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may be decomposed, combined, or partially merged, so the actual execution order may vary depending on the actual situation.

[0033] It should be understood that the terms used in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification and the appended claims, the singular forms "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0034] Embodiments of the present invention provide a visual data processing method, device, and storage medium, which are intended to implement task-based visual data processing to improve the reliability of task execution.

[0035] The following embodiments of the present invention are described in detail with reference to the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments may be combined with each other.

[0036] Please refer to Figure 1 , Figure 1 A flowchart of a visual data processing method provided by an embodiment of the present invention. The visual data processing method can be applied in visual data processing devices or other devices such as computer devices to implement task-based visual data processing to improve the reliability of task execution.

[0037] like Figure 1 As shown, the visual data processing method includes steps S101 to S104.

[0038] S101. Processing acquired visual data through an encoder to obtain high-frequency feature data and low-frequency feature data corresponding to the visual data, wherein the encoder includes an octave convolution module.

[0039] The visual data includes but is not limited to video data, image data, etc. Exemplarily, the encoder includes a plurality of octave convolution modules, for example, Figure 2 As shown, the encoder includes multiple octave convolution modules OctConv and self-calibrated convolution network modules SCC Module (Self-calibrated Convolution). For each octave convolution module, as shown Figure 3 As shown, involving f(X H ;φ H→H ) algorithm unit, f ×2↓ (X H ;φ H→L ) algorithm unit, f ×2↑ (XL ;φ L→H ) algorithm unit, and f(X L ;φ L→L ) algorithm unit, where the visual data X H Respectively through f(X H ;φ H→H ) algorithm unit processing and f ×2↓ (X H ;φ H→L ) algorithm unit performs downsampling processing, and the visual data X L Respectively through f(X L ;φ L→L ) algorithm unit processing and f ×2↑ (X L ;φ L→H ) algorithm unit performs upsampling processing and summarizes f(X H ;φ H→H )The output of the algorithm unit and f ×2↑ (X L ;φ L→H ) The output of the algorithm unit is used to obtain high-frequency feature data Y H , sum up f(X L ;φ L→L )The output of the algorithm unit and f ×2↓ (X H ;φ H→L ) The output of the algorithm unit is used to obtain the low-frequency feature data Y L .

[0040] After being processed by the encoder, the visual data is divided into high-frequency and low-frequency parts, which can reduce the redundancy of low-frequency information in the calculation process, reduce the spatial size of low-frequency feature data, and thus save storage space.

[0041] In some embodiments, as Figure 4 As shown, step S101 may include sub-step S1011 and sub-step S1012.

[0042] S1011, processing the visual data through the octave convolution module to obtain initial high-frequency feature data and initial low-frequency feature data;

[0043] S1012. Input the initial high-frequency feature data and the initial low-frequency feature data into a self-calibration convolutional network module, perform ordinary convolution processing on the initial low-frequency feature data to obtain the low-frequency feature data, and perform self-calibration processing on the initial high-frequency feature data to obtain the high-frequency feature data.

[0044] For example, based on Figure 3 The octave convolution module shown in the figure processes the input visual data to obtain the initial high-frequency feature data X1H and initial low-frequency feature data X1 L , then the obtained initial high-frequency feature data X1 H and initial low-frequency feature data X1 L Input self-calibration convolutional network module, such as Figure 5 As shown, the self-calibrated convolutional network module includes multiple convolutional layers Conv3, where the initial low-frequency feature data X1 L After ordinary convolution processing is performed on the convolution layer Conv3, low-frequency feature data Y is obtained L Initial high-frequency feature data X1 H One branch is processed by the convolution layer Conv3 for ordinary convolution, and the other branch is processed by downsampling and upsampling in sequence, and is combined with the initial high-frequency feature data X1 H The data obtained from the two branches are added together, and then self-calibrated by the Sigmoid function, and then processed by the convolution layer Conv3 for ordinary convolution to obtain high-frequency feature data Y H .

[0045] By performing self-calibration processing on high-frequency components through the self-calibration convolutional network module, the high-frequency components can obtain multi-scale information and improve the accuracy of machine vision tasks.

[0046] For example, the high-frequency feature data Y can be determined based on the current visual image task. H With low-frequency feature data Y L For example, if the current visual image task does not focus on the low-frequency information in the image, the high-frequency feature data Y corresponding to the current visual image task H With low-frequency feature data Y L The ratio is greater than 1, that is, the high-frequency feature data Y H The amount of data is greater than the low-frequency feature data Y L The amount of data.

[0047] For example, taking visual data as image data, assume that the size of the input image is H×W×3, where H and W are the width and height of the image respectively, and 3 represents the three color channels of the image. The proportion of high-frequency and low-frequency component channels in the octave convolution module is actually adjustable. If high-frequency information is more important than low-frequency information, the ratio of the total number of high-frequency and low-frequency component channels can be set to 2:1. At this time, the size of the high-frequency component after passing through the first octave convolution module is The low-frequency component is Where N is the total number of high and low frequency component channels.

[0048] S102. Input the high-frequency feature data into the channel dimension attention module for processing to obtain first high-frequency feature component data and second high-frequency feature component data.

[0049] For example, the channel dimension attention module is as follows Figure 6 As shown in Figure 1, it includes the global pooling layer, the fully connected layer, the activation function ReLU, and the Sigmoid function. The high-frequency feature data Y H As the input of the channel dimension attention module, the output of the channel dimension attention module will be divided into two parts. In order to facilitate the distinction and description, the two parts of the output are referred to as the first high-frequency feature component data Y1 H and the second high frequency feature component data Y2 H .

[0050] In some embodiments, as Figure 7 As shown, step S102 may include sub-step S1021 and sub-step S1022.

[0051] S1021: Input the high-frequency feature data into the channel-dimensional attention module for processing, and determine the attention weight value corresponding to each channel;

[0052] S1022. Perform weighted processing on the high-frequency feature data based on the attention weight value, and divide the weighted processed data into the first high-frequency feature component data and the second high-frequency feature component data.

[0053] Exemplarily, the inputting the high-frequency feature data into the channel-dimensional attention module for processing and determining the attention weight value corresponding to each channel includes:

[0054] The high-frequency feature data is globally averaged pooled by the global pooling layer, and the data after global average pooling is processed by the fully connected layer to obtain the attention weight mapping data corresponding to each channel. The attention weight mapping data is activated by the activation function ReLU and processed again by the fully connected layer. The output data is processed by the Sigmoid function to obtain the attention weight value corresponding to each channel.

[0055] For example, Figure 8 As shown, input high-frequency feature data Y H , one branch for high frequency feature data Y H Perform global average pooling to compress the two-dimensional data into one-dimensional data. Then, use the fully connected layer to map the global feature description of each channel to the attention weight and activate it through the ReLU function. Then use the fully connected layer to restore the dimension. Then, perform Sigmoid activation on the output of the fully connected layer and map the output to between 0 and 1 to obtain the attention weight value corresponding to each channel. Finally, the high-frequency feature data Y of the other branch is converted into the attention weight. HMultiply each channel by the attention weight value for weighted calculation, and divide the obtained data into the first high-frequency feature component data Y1 H and the second high frequency feature component data Y2 H The importance of each channel is learned through the channel-dimensional attention module, and then the useful feature channels are enhanced and the less important feature channels are suppressed according to the channel importance.

[0056] Among them, the first high-frequency characteristic component data Y1 H and the second high frequency feature component data Y2 H It is used to complete different machine vision tasks. Through the channel-dimensional attention module, different machine vision tasks can pay more attention to high-frequency information that is beneficial to the task.

[0057] Still taking the input image listed above as an example, the high-frequency feature data Y H Size Processed by the channel dimension attention module, the first high-frequency feature component data Y1 is obtained H and the second high frequency feature component data Y2 H Size is

[0058] S103. According to the task type of the current visual image task, determine the task target data based on the low-frequency feature data, the first high-frequency feature component data and the second high-frequency feature component data, encode the task target data and obtain a visual image code stream.

[0059] Among them, the task types include but are not limited to machine vision tasks, image reconstruction tasks, etc. The image reconstruction tasks are further divided into image reconstruction tasks of human-machine mixed scenes and image reconstruction tasks of non-human-machine mixed scenes. For tasks of different task types, the first high-frequency feature component data Y1 H , the second high-frequency characteristic component data Y2 H And low-frequency feature data Y L Perform different encoding processes.

[0060] In some embodiments, the first high-frequency feature component data is a high-frequency basic feature, and the second high-frequency feature component data is a high-frequency enhanced feature. The determining of task target data based on the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data according to the task type of the current visual image task includes:

[0061] If the task type of the current visual image task is a machine vision task, determining the low-frequency feature data and the first high-frequency feature component data as the task target data;

[0062] If the task type of the current visual image task is an image reconstruction task, the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data are determined as the task target data.

[0063] For machine vision tasks, the low-frequency feature data Y L and the first high frequency feature component data Y1 H It is used as the task target data and is encoded to complete the machine vision task. The advantage is that the bit rate for completing the machine vision task can be greatly saved when image reconstruction is not required.

[0064] For the image reconstruction task, the low-frequency feature data Y L , first high frequency feature component data Y1 H and the second high frequency feature component data Y2 H As the task target data, it is encoded and processed, and the entire code stream is used to complete the image reconstruction task. The advantage is that it will not have a significant impact on its quality.

[0065] In some embodiments, encoding the task target data to obtain a visual image code stream includes:

[0066] If the task type is an image reconstruction task, determining whether the current application scenario is a human-machine hybrid scenario;

[0067] If it is not a human-machine mixed scene, encoding the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data to obtain a visual image code stream;

[0068] If it is a human-machine mixed scene, the low-frequency feature data and the first high-frequency feature component data are encoded, and the second high-frequency feature component data is encoded, and the two parts of the encoding results are merged to obtain the visual image code stream corresponding to the image reconstruction task.

[0069] For example, for the image reconstruction task, it is determined whether the current application scene is a human-machine mixed scene. If it is not a human-machine mixed scene, the low-frequency feature data Y L , first high frequency feature component data Y1 H and the second high frequency feature component data Y2 H , that is, all features are encoded. If it is a human-machine mixed scene, in addition to the low-frequency feature data Y L , first high frequency feature component data Y1 H Encoding is performed, and the remaining second high frequency feature component data Y2 is encoded. H Encoding is also performed.

[0070] In some embodiments, encoding the task target data to obtain a visual image code stream includes:

[0071] quantizing the task target data to obtain a quantized feature vector;

[0072] The quantized feature vector is processed by a context model, and the task target data is processed by a super priori model. The processing results of the context model and the super priori model are input into the entropy model parameter prediction module for processing to obtain the corresponding mean and variance;

[0073] Encoding is performed according to the quantized eigenvector, the mean, and the variance to obtain the visual image code stream.

[0074] After determining the task target data corresponding to the current visual image task, the task target data is quantized by the quantization module to obtain a quantized feature vector, which is input into the context model and used as the input of the entropy model parameter prediction module after being processed by the context model. In addition, the task target data is also used as the input of the entropy model parameter prediction module after being processed by the super prior model, wherein the super prior model includes multiple components such as a super parameter encoder and a super parameter decoder. The corresponding mean and variance are obtained by the entropy model parameter prediction module, and are input into the corresponding arithmetic encoder for encoding processing to obtain the encoded code stream. For example, Figure 9 As shown, the hyperparameter encoder also includes multiple octave convolution modules OctConv, which convert the low-frequency feature data Y L , first high frequency feature component data Y1 H and the second high frequency feature component data Y2 H As the input of the hyperparameter encoder, the output feature vector Z L and Z H For example, the entropy model parameter prediction module has three convolutional layers Conv, which can obtain the mean and variance of the corresponding Gaussian distribution.

[0075] Specifically, if Figure 10 As shown, at the encoding end, after the input image is processed by the encoder and the channel dimension attention module, {Y L , Y1 H , Y2 H}After being processed by the hyperparameter encoder, quantization module Q, and arithmetic encoder AE, the corresponding coded data is obtained, which is transmitted to the arithmetic decoder AD and processed by the arithmetic decoder AD and the hyperparameter decoder as part of the input of the entropy model parameter prediction module EP. L , Y1 H , Y2 H} is processed by the quantization module Q and the context model CTX as another part of the input of the entropy model parameter prediction module EP. The entropy model parameter prediction module EP outputs the corresponding Gaussian distribution mean and variance, which are the same as {Y L , Y1 H , Y2 H}Input the corresponding arithmetic encoder AE to obtain the bit stream.

[0076] During the encoding process, since the encoding process is carried out in order of spatial position, that is, processing row by row from left to right, when encoding the features of the current m-th spatial position, the input of the context model CTX is the features before the m-th position (that is, the subscript is less than m). The context model CTX masks the features of m and after the m-th position. Then, the high-frequency features obtained by the context model CTX are combined with the parameters of the hyperparameter decoder before the m-th position as the input of the entropy model parameter prediction module EP to obtain the mean and variance of the features of the current m-th position. Finally, the obtained mean and variance are input to the arithmetic encoder AE to perform arithmetic encoding on the features.

[0077] For example, the code stream structure is shown in Table 1:

[0078] Table 1

[0079] Field meaning Occupied bits Encoding machine_task_num Number of machine vision tasks 8-bit Fixed-length encoding image_recon_flag Whether to reconstruct the image 1 person Fixed-length encoding channel_num Total number of channels 8-bit Fixed-length encoding begin_channel_idx Machine vision task starting channel 8-bit Fixed-length encoding end_channel_idx Machine vision task termination channel 8-bit Fixed-length encoding payload Data content variable length variable-length encoding Check digit Data integrity check 4 digits Fixed-length encoding

[0080] The code stream structure can be organized according to the common coding method of the international specification of the video coding standard. For example, the fixed-length coded syntax elements in the above table can be encoded as unsigned integer fixed-length codes u(n), where n represents the number of occupied bits, such as the machine_task_num field can be represented as u(8), and the image_recon_flag field can be represented as u(1). The variable-length coded syntax elements can be encoded as signed integer zero-order Golomb variable-length codes se(v).

[0081] S104: Transmit the visual image code stream to a decoding end for decoding to obtain decoded data, and complete the current visual image task according to the decoded data.

[0082] At the decoding end, when completing the machine vision task, the low-frequency feature data Y is decoded from the entropy model parameter prediction module L , first high frequency feature component data Y1 H , then merge and reconstruct features; when the image reconstruction task is completed, the second high-frequency feature component data Y2 is decoded from the entropy model parameter prediction module H , and the low-frequency feature data Y L , first high frequency feature component data Y1 H , the second high-frequency characteristic component data Y2 HThe three code streams are merged and decoded.

[0083] In some embodiments, transmitting the visual image code stream to a decoding end for decoding, obtaining decoded data, and completing the current visual image task according to the decoded data includes:

[0084] The visual image code stream is transmitted to a decoding end, processed by an entropy model parameter prediction module, and then input into an arithmetic decoder for processing to obtain the decoded data, wherein the decoded data includes the low-frequency feature data and the first high-frequency feature component data;

[0085] Combining the low-frequency feature data and the first high-frequency feature component data and inputting the combined data into a multi-scale feature reconstruction network to perform multi-scale feature fusion processing to obtain machine vision task data;

[0086] The machine vision task is completed according to the machine vision task data.

[0087] For example, Figure 11 As shown in Figure 2, the multi-scale feature reconstruction network includes modules such as FT1 (Feature Transform), FT2, downsampling convolution layer (↓2), and multiple convolution layers Conv. Figure 12 As shown in the figure, the FT1 and FT2 modules include RB (Residule Block), activation function Leaky Relu, upsampling convolution layer (↑2), etc. RB includes multiple convolution layers Conv, activation function Leaky Relu, etc.

[0088] Specifically, if Figure 13 As shown in the figure, at the decoding end, the decoding process is also carried out in the order of spatial positions. When encoding the features of the current m-th spatial position, the input of the context model CTX is the reconstructed values ​​of all features before the m-th position (that is, the subscript is less than m) that have been decoded. The m-th position and subsequent positions are padded with zeros as the input of the context model CTX to obtain high-frequency features. The high-frequency features are combined with the decoding parameters of the hyperparameter decoder of the first m spatial positions as the input of the entropy model parameter prediction module EP to obtain the mean and variance of the features at the current position. Finally, the mean and variance are input to the arithmetic decoder AD for decoding.

[0089] For machine vision tasks such as Figure 13 and Figure 14 As shown, the low-frequency feature data Y L , first high frequency feature component data Y1 H As the input of the multi-scale feature reconstruction network, the two inputs are spliced ​​by channel. For example, taking the input image listed above as an example, the low-frequency feature data Y L, first high frequency feature component data Y1 H Size is As the input of the multi-scale feature reconstruction network, the size of the intermediate feature vector obtained after splicing is Feature reconstruction is then performed to obtain the P2, P3, P4, and P5 machine vision task data inputs for the backend machine vision network. For example, the machine vision network is the second half of the Mask RCNN X-101FPN. Based on the P2, P3, P4, and P5 machine vision task data, the machine vision task is completed.

[0090] From the above operation process, it can be seen that the machine vision task is completed based on the multi-scale feature reconstruction network, skipping the decoding process, thereby effectively reducing the computational complexity of the machine vision task.

[0091] In some embodiments, transmitting the visual image code stream to a decoding end for decoding, obtaining decoded data, and completing the current visual image task according to the decoded data includes:

[0092] The visual image code stream is transmitted to a decoding end, processed by an entropy model parameter prediction module, and then input into an arithmetic decoder for processing to obtain the decoded data, wherein the decoded data includes the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data;

[0093] Processing the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data through a decoder to obtain image reconstruction task data;

[0094] The image reconstruction task is completed according to the image reconstruction task data.

[0095] like Figure 13 and Figure 14 As shown, for the image reconstruction task, the first high-frequency feature component data Y1 H and the second high frequency feature component data Y2 H For example, still taking the input image listed above as an example, the first high-frequency feature component data Y1 H , the second high-frequency characteristic component data Y2 H Size is The high-frequency feature vector size obtained after connection is The merged high-frequency feature vector and low-frequency feature vector are then input into a decoder composed of an octave convolution module OctConv and a self-calibration convolutional network module SCC Module. Based on the image reconstruction task data, the decoder is used to complete the image reconstruction task.

[0096] The encoder, hyperparameter encoder, decoder, and hyperparameter decoder structures used are all optimized using the OctConv convolution module, which reduces the spatial size of low-frequency components, can reduce bit rate, and reduce computational complexity. For example, the encoder, hyperparameter encoder, decoder, and hyperparameter decoder structure parameters are shown in Table 2:

[0097] Table 2

[0098]

[0099] For example, the entire process of executing each task is as follows: Figure 15 As shown, first determine the task type based on actual needs or application scenarios, where the task types mainly include three types: machine vision tasks, non-human-machine mixed image reconstruction tasks, and human-machine mixed image reconstruction tasks. In this solution, first obtain the high- and low-frequency feature space through the encoder to determine whether the image reconstruction task needs to be completed. If the image reconstruction task does not need to be completed, select some channels based on the amount of information required for the machine vision task, and then complete each machine vision task through arithmetic encoding and decoding and multi-scale feature reconstruction network. If the image reconstruction task needs to be completed, all channels need to be selected for arithmetic encoding, and then determine whether it is a human-machine mixed application scenario. If it is not a human-machine mixed scenario, decode the entire code stream arithmetic decoding to achieve human eye viewing; if it is a human-machine mixed scenario, first arithmetically decode the feature channels corresponding to the machine vision task, and finally arithmetically decode the remaining channels to complete the image reconstruction task.

[0100] The number of channels for a task can be adjusted based on the accuracy requirements and complexity of different tasks. The higher the accuracy requirements and the more complex the task, the more channels are allocated. However, in human-machine mixed scenarios, the number of machine vision task channels should not exceed the corresponding proportion of the total number of channels (for example, 3 / 4). If the number of machine vision task channels exceeds 3 / 4 of the total number of channels, it may affect the accuracy of the image reconstruction task.

[0101] For example, in the product segmentation scenario in the field of industrial automation, the machine vision task of instance segmentation needs to be completed. The specific operation process is as follows:

[0102] S1: Input the product image to be segmented, and use an encoder composed of four downsampling octave convolution modules and a self-calibrated convolutional network module to obtain high-frequency features and low-frequency features. The ratio of the number of high-frequency feature channels to low-frequency feature channels is selected as 1:1.

[0103] S2: Use the channel dimension attention module to enhance the importance of channels related to the instance segmentation task and separate them from all high-frequency features to obtain high-frequency base layer features and high-frequency enhancement layer features.

[0104] S3: Use the entropy model parameter prediction module, context model and arithmetic encoder to encode low-frequency features and high-frequency base layer features without encoding high-frequency enhancement layer features.

[0105] S4: On the intelligent analysis side, high-frequency base layer features and low-frequency features are decoded and mapped to the inputs of P2, P3, P4, and P5 of the back-end machine vision network Mask RCNN X-101FPN through a multi-scale feature reconstruction network. This skips the image decoding process, saving bitrate and reducing computational effort.

[0106] For example, in actual applications, encoding-related components including encoder, channel dimension attention module, hyperparameter encoder, hyperparameter decoder, context model, entropy model parameter prediction module, arithmetic encoder, arithmetic decoder, etc. are deployed at the image acquisition end, and decoding-related components including hyperparameter decoder, context model, entropy model parameter prediction module, arithmetic decoder and multi-scale feature reconstruction module are deployed at the intelligent analysis end that completes the segmentation.

[0107] For example, in an intelligent surveillance scenario, two tasks need to be completed: target detection and human viewing. For a mixed human-machine scenario, the specific operation process is as follows:

[0108] S1: Input the original image and divide the feature space into high-frequency and low-frequency parts through the octave convolution module, reducing the spatial size of the low-frequency features to save bit rate, and using the self-calibration convolutional network module to enable the high-frequency features to obtain multi-scale information, thereby improving the accuracy of completing machine vision tasks.

[0109] S2: The high-frequency components required to complete the target detection task are separated through the channel dimension attention module. Since missed detection is not desired in the monitoring scenario, the proportion of high-frequency base layer features can be selected as 1 / 2 or a larger value in the embodiment.

[0110] S3: Encode high-frequency base layer features, high-frequency enhancement layer features, and low-frequency features.

[0111] S4: Decode the high-frequency base layer features and low-frequency features and input them into the multi-scale feature reconstruction network to obtain the input of the machine vision network FasterRCNN X-101FPN; and decode the high-frequency enhancement layer features, merge them with the decoded high-frequency base layer features and low-frequency features and input them into the decoder to complete the image reconstruction task.

[0112] For example, in this solution, encoding components and decoding components, including the decoder, can be deployed on the video capture end and the user end, respectively. Based on the feedback from the user end, encoding of a portion of the bitstream or the entire bitstream can be selected. If only target detection is required to detect anomalies in the surveillance scene, the signal is transmitted to the encoder, which encodes the basic bitstream and transmits it to the user end. The user end directly maps the features to the backend machine vision network to complete the target detection task. If image viewing is required, the entire bitstream is encoded.

[0113] In the above embodiment, the acquired visual data is processed by the encoder to obtain high-frequency feature data and low-frequency feature data corresponding to the visual data, and the high-frequency feature data is input into the channel dimension attention module for processing to obtain first high-frequency feature component data and second high-frequency feature component data. According to the task type of the current visual image task, the task target data is determined based on the low-frequency feature data, the first high-frequency feature component data and the second high-frequency feature component data, the task target data is encoded and processed to obtain a visual image code stream, the visual image code stream is transmitted to the decoding end for decoding, the decoded data is obtained, and the current visual image task is completed according to the decoded data. That is, for various tasks, the tasks are completed by performing different data processing methods on the visual data, thereby improving the reliability of task execution.

[0114] The present invention also provides a visual data processing device. Figure 16 , Figure 16 It is a schematic block diagram of a visual data processing device provided in one embodiment of the present application.

[0115] like Figure 16 As shown, the visual data processing device 200 may include a processor 210 and a memory 220, wherein the processor 210 and the memory 220 are connected via a bus, such as an I2C (Inter-integrated Circuit) bus.

[0116] Specifically, the processor 210 may be a micro-controller unit (MCU), a central processing unit (CPU), or a digital signal processor (DSP).

[0117] Specifically, the memory 220 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a mobile hard disk, etc. The memory 220 stores various computer programs for execution by the processor 210 .

[0118] The processor 210 is configured to run a computer program stored in the memory and implement the following steps when executing the computer program:

[0119] Processing the acquired visual data through an encoder to obtain high-frequency feature data and low-frequency feature data corresponding to the visual data, wherein the encoder includes an octave convolution module;

[0120] Inputting the high-frequency feature data into the channel dimension attention module for processing to obtain first high-frequency feature component data and second high-frequency feature component data;

[0121] According to the task type of the current visual image task, determining task target data based on the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data, encoding the task target data to obtain a visual image code stream;

[0122] The visual image code stream is transmitted to a decoding end for decoding to obtain decoded data, and the current visual image task is completed according to the decoded data.

[0123] In some embodiments, when the processor 210 implements inputting the high-frequency feature data into the channel dimension attention module for processing to obtain the first high-frequency feature component data and the second high-frequency feature component data, it is configured to implement:

[0124] Input the high-frequency feature data into the channel dimension attention module for processing, and determine the attention weight value corresponding to each channel;

[0125] The high-frequency feature data is weighted based on the attention weight value, and the weighted data is divided into the first high-frequency feature component data and the second high-frequency feature component data.

[0126] In some embodiments, the channel-dimensional attention module includes a global pooling layer, a fully connected layer, an activation function ReLU, and a Sigmoid function. When the processor 210 inputs the high-frequency feature data into the channel-dimensional attention module for processing and determines the attention weight value corresponding to each channel, it is used to implement:

[0127] The high-frequency feature data is globally averaged pooled by the global pooling layer, and the data after global average pooling is processed by the fully connected layer to obtain the attention weight mapping data corresponding to each channel. The attention weight mapping data is activated by the activation function ReLU and processed again by the fully connected layer. The output data is processed by the Sigmoid function to obtain the attention weight value corresponding to each channel.

[0128] In some embodiments, when the processor 210 encodes the task target data to obtain a visual image code stream, it is configured to implement:

[0129] quantizing the task target data to obtain a quantized feature vector;

[0130] The quantized feature vector is processed by a context model, and the task target data is processed by a super priori model. The processing results of the context model and the super priori model are input into the entropy model parameter prediction module for processing to obtain the corresponding mean and variance;

[0131] Encoding is performed according to the quantized eigenvector, the mean, and the variance to obtain the visual image code stream.

[0132] In some embodiments, the task type includes a machine vision task and an image reconstruction task, the first high-frequency feature component data is a high-frequency basic feature, and the second high-frequency feature component data is a high-frequency enhanced feature; when the processor 210 determines the task target data based on the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data according to the task type of the current visual image task, it is configured to implement:

[0133] If the task type of the current visual image task is a machine vision task, determining the low-frequency feature data and the first high-frequency feature component data as the task target data;

[0134] If the task type of the current visual image task is an image reconstruction task, the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data are determined as the task target data.

[0135] In some embodiments, if the task type of the current visual image task is a machine vision task, the processor 210, when implementing the transmission of the visual image code stream to the decoding end for decoding, obtaining decoded data, and completing the current visual image task according to the decoded data, is configured to implement:

[0136] The visual image code stream is transmitted to a decoding end, processed by an entropy model parameter prediction module, and then input into an arithmetic decoder for processing to obtain the decoded data, wherein the decoded data includes the low-frequency feature data and the first high-frequency feature component data;

[0137] Combining the low-frequency feature data and the first high-frequency feature component data and inputting the combined data into a multi-scale feature reconstruction network to perform multi-scale feature fusion processing to obtain machine vision task data;

[0138] The machine vision task is completed according to the machine vision task data.

[0139] In some embodiments, when the processor 210 encodes the task target data to obtain a visual image code stream, it is configured to implement:

[0140] If the task type is an image reconstruction task, determining whether the current application scenario is a human-machine hybrid scenario;

[0141] If it is not a human-machine mixed scene, encoding the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data to obtain the visual image code stream;

[0142] If it is a human-machine mixed scene, the low-frequency feature data and the first high-frequency feature component data are encoded, and the second high-frequency feature component data is encoded, and the two parts of the encoding results are merged to obtain the visual image code stream corresponding to the image reconstruction task.

[0143] In some embodiments, if the task type of the current visual image task is an image reconstruction task, the processor 210, when implementing the transmitting of the visual image code stream to the decoding end for decoding, obtaining decoded data, and completing the current visual image task according to the decoded data, is configured to implement:

[0144] The visual image code stream is transmitted to a decoding end, processed by an entropy model parameter prediction module, and then input into an arithmetic decoder for processing to obtain the decoded data, wherein the decoded data includes the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data;

[0145] Processing the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data through a decoder to obtain image reconstruction task data;

[0146] The image reconstruction task is completed according to the image reconstruction task data.

[0147] In some embodiments, when the processor 210 processes the acquired visual data through the encoder to obtain high-frequency feature data and low-frequency feature data corresponding to the visual data, it is configured to implement:

[0148] Processing the visual data through the octave convolution module to obtain initial high-frequency feature data and initial low-frequency feature data;

[0149] The initial high-frequency feature data and the initial low-frequency feature data are input into a self-calibration convolutional network module, the initial low-frequency feature data is subjected to ordinary convolution processing to obtain the low-frequency feature data, and the initial high-frequency feature data is subjected to self-calibration processing to obtain the high-frequency feature data.

[0150] The visual data processing device 200 can execute the visual data processing method provided in the embodiment of the present application, and therefore can achieve the beneficial effects that can be achieved by the visual data processing method provided in the embodiment of the present application. Please refer to the previous embodiment for details and will not be repeated here.

[0151] An embodiment of the present application also provides a storage medium for computer-readable storage, wherein the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of any visual data processing method provided in the description of the embodiment of the present application.

[0152] The storage medium may be an internal storage unit of the terminal device described in the aforementioned embodiment, such as a hard disk or memory of the terminal device. The storage medium may also be an external storage device of the terminal device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash memory card, etc. equipped on the terminal device.

[0153] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In a hardware embodiment, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0154] It should be understood that the term "and / or" used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations. It should be noted that, in this article, the terms "include", "comprise" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or system that includes a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article or system. In the absence of further restrictions, an element defined by the sentence "including a..." does not exclude the presence of other identical elements in the process, method, article or system that includes the element.

[0155] The serial numbers of the embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments. The above are only specific embodiments of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application shall be based on the scope of protection of the claims.

Claims

1. A method for processing visual data, the method comprising: Processing the acquired visual data through an encoder to obtain high-frequency feature data and low-frequency feature data corresponding to the visual data, wherein the encoder includes an octave convolution module; Inputting the high-frequency feature data into a channel dimension attention module for processing to obtain first high-frequency feature component data and second high-frequency feature component data; According to the task type of the current visual image task, determining task target data based on the low-frequency feature data, the first high-frequency feature component data and the second high-frequency feature component data, encoding the task target data to obtain a visual image code stream; The visual image code stream is transmitted to a decoding end for decoding to obtain decoded data, and the current visual image task is completed according to the decoded data.

2. The method according to claim 1, characterized in that: The inputting the high-frequency feature data into the channel dimension attention module for processing to obtain the first high-frequency feature component data and the second high-frequency feature component data comprises: Input the high-frequency feature data into the channel dimension attention module for processing, and determine the attention weight value corresponding to each channel; The high-frequency feature data is weighted based on the attention weight value, and the weighted data is divided into the first high-frequency feature component data and the second high-frequency feature component data.

3. The method according to claim 2, characterized in that The channel dimension attention module includes a global pooling layer, a fully connected layer, an activation function ReLU, and a Sigmoid function. The high-frequency feature data is input into the channel dimension attention module for processing to determine the attention weight value corresponding to each channel, including: The high-frequency feature data is subjected to global average pooling processing through the global pooling layer. The data after global average pooling processing is processed by the fully connected layer to obtain attention weight mapping data corresponding to each channel. The attention weight mapping data is activated by the activation function ReLU and processed again by the fully connected layer. The output data is processed by the Sigmoid function to obtain the attention weight value corresponding to each channel.

4. The method according to claim 1, characterized in that: The encoding process of the task target data to obtain a visual image code stream includes: Quantizing the task target data to obtain a quantized feature vector; The quantized feature vector is processed by a context model, and the task target data is processed by a super priori model, and the processing results of the context model and the super priori model are input into the entropy model parameter prediction module for processing to obtain the corresponding mean and variance; Encoding processing is performed according to the quantized eigenvector, the mean value and the variance to obtain the visual image code stream.

5. The method according to claim 1, characterized in that The task types include machine vision tasks and image reconstruction tasks, the first high-frequency feature component data is a high-frequency basic feature, and the second high-frequency feature component data is a high-frequency enhanced feature; the task target data is determined based on the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data according to the task type of the current visual image task, including: If the task type of the current visual image task is a machine vision task, determining the low-frequency feature data and the first high-frequency feature component data as the task target data; If the task type of the current visual image task is an image reconstruction task, the low-frequency feature data, the first high-frequency feature component data and the second high-frequency feature component data are determined as the task target data.

6. The method according to claim 5, characterized in that If the task type of the current visual image task is a machine vision task, transmitting the visual image code stream to a decoding end for decoding, obtaining decoded data, and completing the current visual image task according to the decoded data includes: The visual image code stream is transmitted to a decoding end, processed by an entropy model parameter prediction module, and then input into an arithmetic decoder for processing to obtain the decoded data, wherein the decoded data includes the low-frequency feature data and the first high-frequency feature component data; The low-frequency feature data and the first high-frequency feature component data are combined and input into a multi-scale feature reconstruction network to perform multi-scale feature fusion processing to obtain machine vision task data; The machine vision task is completed according to the machine vision task data.

7. The method according to claim 5, characterized in that The encoding process of the task target data to obtain a visual image code stream includes: If the task type is an image reconstruction task, determining whether the current application scenario is a human-machine hybrid scenario; If it is not a human-machine mixed scene, encoding the low-frequency feature data, the first high-frequency feature component data and the second high-frequency feature component data to obtain the visual image code stream; If it is a human-machine mixed scene, the low-frequency feature data and the first high-frequency feature component data are encoded, and the second high-frequency feature component data is encoded, and the two parts of the encoding results are merged to obtain the visual image code stream corresponding to the image reconstruction task.

8. The method according to claim 7, characterized in that If the task type of the current visual image task is an image reconstruction task, transmitting the visual image code stream to a decoding end for decoding, obtaining decoded data, and completing the current visual image task according to the decoded data includes: The visual image code stream is transmitted to a decoding end, processed by an entropy model parameter prediction module, and then input into an arithmetic decoder for processing to obtain the decoded data, wherein the decoded data includes the low-frequency feature data, the first high-frequency feature component data, and the second high-frequency feature component data; Processing the low-frequency feature data, the first high-frequency feature component data and the second high-frequency feature component data through a decoder to obtain image reconstruction task data; The image reconstruction task is completed according to the image reconstruction task data.

9. The method according to claim 1, characterized in that: The encoder processes the acquired visual data to obtain high-frequency feature data and low-frequency feature data corresponding to the visual data, including: Processing the visual data through the octave convolution module to obtain initial high-frequency feature data and initial low-frequency feature data; The initial high-frequency feature data and the initial low-frequency feature data are input into a self-calibration convolutional network module, the initial low-frequency feature data is subjected to ordinary convolution processing to obtain the low-frequency feature data, and the initial high-frequency feature data is subjected to self-calibration processing to obtain the high-frequency feature data.

10. A visual data processing device, characterized in that: The visual data processing device comprises: A processor, a memory, a computer program stored in the memory and executable by the processor, and a data bus for realizing connection and communication between the processor and the memory, wherein when the computer program is executed by the processor, the steps of the visual data processing method as described in any one of claims 1 to 9 are realized.

11. A storage medium for computer-readable storage, characterized in that: The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the visual data processing method as described in any one of claims 1 to 9.