Image Recognition Method, Apparatus and Computer-Readable Storage Medium
The method enhances light field image distortion detection by fusing sub-aperture features and generating distortion scores, addressing limitations in existing methods and improving accuracy and reliability.
Patent Information
- Application Number
- CN202110610183.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-06-01
AI Technical Summary
In the prior art, when the light field image is distorted, the method of degrading characteristics of the angle consistency of the extreme plane image based on the light field image is only applicable to typical distortion, making it difficult to accurately identify the low-distorted light field image, and the degree of distortion cannot be determined, and the accuracy and reliability are low.
By receiving the light field image, the sub-aperture feature set is extracted, and the correlation parameters and distortion parameters are generated based on the correlation and recognition degree of the sub-aperture fusion features under different object types, the target distortion score of the light field image is determined, and the deep learning model and the global context perception module are used for feature fusion and recognition.
Accurate recognition of light field images of any degree of distortion is achieved, precise determination of the degree of distortion of light field images, and the accuracy and reliability of image recognition are improved.
Smart Images

Figure CN113822115B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to an image recognition method, apparatus, and computer-readable storage medium. Background Art
[0002] With the development of optoelectronic technology, light field imaging technology can obtain objects at multiple angles of a single object or image, such as the position and direction information of light during propagation. Compared with traditional imaging, light field imaging contains richer visual information. However, in the process of processing the light field information, light field imaging may cause distortion of the light field image; in order to identify whether there is distortion in the light field image, related technologies calculate the distribution of the gradient direction map of the epipolar plane image (EPI) according to the spatial quality of the light field image to determine the local angular consistency degradation characteristics, so as to identify the distorted light field image.
[0003] In the process of researching and practicing the prior art, the inventors of the present invention found that when identifying whether there is distortion in the light field image in the prior art, it is determined according to the degradation characteristics of the angular consistency of the epipolar plane image of the light field image. This method identifies the distortion of the image through limited image features, which is only applicable to typical distortions of light field images, and it is difficult to accurately identify low-distortion light field images, and it is impossible to determine the degree of distortion of the light field image, resulting in low reliability. Summary of the Invention
[0004] Embodiments of the present application provide an image recognition method, apparatus, and computer-readable storage medium, which can improve the accuracy of image recognition.
[0005] Embodiments of the present application provide an image recognition method, including:
[0006] Receiving a light field image to be recognized, where the light field image includes light sub-images at different light field angles;
[0007] Performing feature extraction on each light sub-image to obtain a sub-aperture feature set corresponding to each light sub-image;
[0008] Fusing multiple sub-aperture feature sets according to the object type to obtain sub-aperture fusion features under different object types;
[0009] Generating corresponding correlation parameters according to the correlation degree between the sub-aperture fusion features under different object types, and generating distortion parameters according to the recognition degree of the sub-aperture fusion features under different object types, where the correlation degree is generated by the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated by the difference degree between the sub-aperture fusion features of different object types;
[0010] Determining a target distortion score of the light field image according to the correlation parameters and the distortion parameters.
[0011] Correspondingly, an embodiment of the present application provides an image recognition device, including:
[0012] A receiving unit, configured to receive a light field image to be recognized, where the light field image includes sub-light field images at different light field angles;
[0013] An extraction unit, configured to perform feature extraction on each sub-light field image to obtain a set of sub-aperture features corresponding to each sub-light field image;
[0014] A fusion unit, configured to fuse the sets of sub-aperture features according to object types to obtain sub-aperture fusion features under different object types;
[0015] A generation unit, configured to generate corresponding correlation parameters according to the correlation degree between the sub-aperture fusion features under different object types, and generate distortion parameters according to the recognition degree of the sub-aperture fusion features under different object types, where the correlation degree is generated by the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated by the difference degree between the sub-aperture fusion features of different object types;
[0016] A determination unit, configured to determine a target distortion score of the light field image according to the correlation parameters and the distortion parameters.
[0017] In some embodiments, the fusion unit includes:
[0018] An identification subunit, configured to identify the object type of each sub-aperture feature in each set of sub-aperture features;
[0019] An acquisition subunit, configured to acquire the color channels of the sub-aperture features corresponding to each object type;
[0020] A fusion subunit, configured to fuse the sub-aperture features of the same object type between the sets of sub-aperture features according to the screening rules of color channels to obtain sub-aperture fusion features under different object types.
[0021] In some embodiments, the fusion subunit is further configured to:
[0022] Fuse the sub-aperture features with the same color channels between the sets of sub-aperture features to obtain first sub-aperture fusion features corresponding to different color channels;
[0023] Filter the first sub-aperture fusion features to obtain filtered second sub-aperture fusion features;
[0024] Fuse the second sub-aperture fusion features of different color channels to obtain sub-aperture fusion features under different object types.
[0025] In some embodiments, the generating unit is further configured to:
[0026] Fuse the sub-aperture features with the same color channel among the sub-aperture feature sets to obtain first sub-aperture fusion features corresponding to different color channels;
[0027] Filter the first sub-aperture fusion features to obtain filtered second sub-aperture fusion features;
[0028] Fuse the second sub-aperture fusion features of different color channels to obtain sub-aperture fusion features under different object types.
[0029] In some embodiments, the generating unit is further configured to:
[0030] Input the sub-aperture fusion features of different object types into the global context awareness module in the trained target model;
[0031] Output association parameters through the global context awareness module, where the association parameters are generated by the global context awareness module according to the association degree between the sub-aperture fusion features of different object types.
[0032] In some embodiments, the device further includes: a training unit, configured to:
[0033] Obtain a sample light field image set, where the sample light field image set includes a plurality of distorted sample light field images and a sample distortion score corresponding to each sample light field image;
[0034] Jointly pre-train a preset model according to the sample light field images and the corresponding sample distortion scores to obtain a pre-trained model;
[0035] Obtain a real data set, where the real data set includes light field images with different compression ratios and a real distortion score corresponding to each light field image;
[0036] Jointly train the pre-trained model according to the light field images with different compression ratios and the corresponding real distortion scores to obtain a target model;
[0037] Then the receiving unit is further configured to receive the light field image to be recognized through the target model.
[0038] In some embodiments, the training unit is further configured to:
[0039] Input the light field images with different compression ratios into the pre-trained model to obtain corresponding predicted distortion scores;
[0040] Obtain the difference value between the predicted distortion score and the real distortion score;
[0041] Iteratively train the network parameters of the pre-trained model based on the difference value until the difference value converges, and obtain the trained target model.
[0042] In some embodiments, the apparatus further includes: a processing unit, configured to:
[0043] Obtain the number of sub-aperture features included in the sub-aperture fusion feature corresponding to each object type;
[0044] Filter the sub-aperture fusion features with the number of sub-aperture features less than a preset feature number threshold to obtain filtered sub-aperture fusion features;
[0045] Perform downsampling processing on the filtered sub-aperture fusion features to obtain target sub-aperture fusion features;
[0046] Then, the generating unit is further configured to generate corresponding association parameters according to the association degree between the target sub-aperture fusion features of different object types, and generate distortion parameters according to the recognition degree of the target sub-aperture fusion features of different object types.
[0047] Correspondingly, the present application further provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps in any one of the image recognition methods provided in the embodiments of the present application are implemented.
[0048] In addition, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in any one of the image recognition methods provided in the embodiments of the present application are implemented.
[0049] In addition, an embodiment of the present application further provides a computer program, the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in any one of the image recognition methods provided in the embodiments of the present application.
[0050] Embodiments of the present application can receive a light field image to be recognized, where the light field image includes light sub-images at different light field angles; extract features from each light sub-image to obtain a set of sub-aperture features corresponding to each light sub-image; fuse the multiple sets of sub-aperture features according to the object type to obtain sub-aperture fusion features under different object types; generate corresponding correlation parameters based on the correlation degrees between the sub-aperture fusion features under different object types, and generate distortion parameters based on the recognition degrees of the sub-aperture fusion features under different object types, where the correlation degree is generated from the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated from the difference degree between the sub-aperture fusion features under different object types; determine the target distortion score of the light field image according to the correlation parameters and the distortion parameters. Thus, by fusing the sub-aperture features of the same object type between the light sub-images in the light field image to obtain sub-aperture fusion features, determining the correlation coefficient according to the correlation features between the sub-aperture fusion features, and determining the distortion coefficient according to the sub-aperture fusion features, and then determining the target quality score of the light field image, in this way, realizing distortion recognition by combining the global sub-aperture features in the light field image, accurately identifying the light field image with any degree of distortion, accurately determining the degree of distortion of the light field image, accurately identifying the degree of image distortion, improving the accuracy of image recognition, and having reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0052] Figure 1 is a schematic diagram of the scenario of the image recognition method provided by the embodiments of the present application;
[0053] Figure 2 is a schematic flowchart of the image recognition method provided by the embodiments of the present application;
[0054] Figure 3 is another schematic flowchart of the image recognition method provided by the embodiments of the present application;
[0055] Figure 4 is a schematic structural diagram of the preset model provided by the embodiments of the present application;
[0056] Figure 5 is a schematic diagram of the sub-aperture feature fusion process provided by the embodiments of the present application;
[0057] Figure 6 is a schematic structural diagram of the image recognition device provided by the embodiments of the present application;
[0058] Figure 7 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. Specific embodiments
[0059] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts belong to the scope of protection of the present invention.
[0060] An embodiment of the present application provides an image recognition method, device, and computer-readable storage medium. Specifically, the embodiment of the present application will be described from the perspective of an image recognition device. The image recognition device can be specifically integrated in a computer device, and the computer device can be a server or a terminal device, etc. Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and the present application does not limit this.
[0061] For example, refer to Figure 1 , which is a schematic diagram of the scenario of the image recognition method provided by the embodiment of the present application. The scenario includes a terminal 10 and a server 20, and the terminal 10 and the server 20 are wirelessly communicatively connected to implement data interaction.
[0062] Among them, the user can send a light field image to the server 20 through the terminal 10; so that the server 20 processes the light field image to determine the target distortion score of the light field image, and returns the target distortion score corresponding to the light field image to the terminal 10; the user can learn about the distortion degree of the current light field image according to the target distortion score received by the terminal 10.
[0063] Among them, the server 20 is used to receive the light field image to be recognized, and the light field image contains light sub-images with different light field angles; extract features from each light sub-image to obtain a set of sub-aperture features corresponding to each light sub-image; fuse the multiple sets of sub-aperture features according to the object type to obtain sub-aperture fusion features under different object types; generate corresponding correlation parameters according to the correlation degree between the sub-aperture fusion features under different object types, and generate distortion parameters according to the recognition degree of the sub-aperture fusion features under different object types, where the correlation degree is generated by the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated by the difference degree between the sub-aperture fusion features of different object types; determine the target distortion score of the light field image according to the correlation parameters and the distortion parameters. Furthermore, return the target distortion score of the light field image to the terminal 10.
[0064] Among them, the image recognition process may include processing methods such as feature extraction, feature fusion, parameter generation, and score determination.
[0065] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.
[0066] In the embodiments of the present application, the description will be made from the perspective of the image recognition device, and the image recognition device may be specifically integrated in a computer device such as a terminal or a server. See Figure 2 , Figure 2 which is a schematic flowchart of the steps of an image recognition method provided by the embodiments of the present application. When the processor on the terminal or the server executes the program corresponding to the image recognition method, the specific process of the image recognition method is as follows:
[0067] 101. Receive the light field image to be recognized.
[0068] Among them, the light field can represent the distribution of light rays in a space, which can record the intensity and direction of the light rays, covering all the information of the light rays during propagation. For example, in practical applications, the light field is a four-dimensional light radiation field that simultaneously contains position and direction information in space and can be represented by a four-dimensional parameterization, that is, the light field contains two-dimensional position information and two-dimensional direction information in space.
[0069] Among them, the light field image may be an image that describes the light ray distribution information in space. Through the light field image, the intuitive information of the same thing in different positions and different angles in space can be recorded, that is, the light field image contains multiple light sub-images with different light field angles, and each light sub-image may be a sub-image that records the corresponding position and light ray angle.
[0070] When image information processing is performed on a light field image, such as during compression, transmission, and rendering of the light field image, distortion of the light field image may occur. In order to identify the distortion of each light field image, a light field image to be identified may be received for further light field image recognition. When performing distortion recognition on a light field image, the degree of distortion of the light field image may be determined directly based on feature information in the light field image, or the light field image may be distorted using a trained model, which is not limited here.
[0071] In order to improve the efficiency of distortion recognition of any light field image, the embodiment of the present application establishes a target model through deep learning, and the target model is used to perform distortion recognition on the light field image. By recognizing the light field image through the target model, the efficiency of distortion recognition on the light field image can be improved.
[0072] In some embodiments, before the step of “receiving a light field image to be identified”, the following steps are included:
[0073] (1) obtaining a sample light field image set, where the sample light field image set includes a plurality of distorted sample light field images and a sample distortion score corresponding to each sample light field image;
[0074] The sample light field image set may be a screened or processed sample set, which includes a plurality of light field images with different distortion levels.
[0075] The light field image may be a light field image that has been processed with image information or an image that has not been processed with image information (i.e., an original light field image). For example, a light field image is processed with image information in a compressed manner to obtain a compressed light field image, which may produce distortion to varying degrees; for another example, an image that has not been processed with image information usually does not produce distortion.
[0076] The sample distortion score may be a rough score of the degree of distortion of the light field image, which is used to roughly indicate the degree of distortion of the light field image. It should be noted that different image information processing methods may cause different degrees of distortion in the light field image, and corresponding sample distortion scores may be set for light field images with different degrees of distortion. For example, when a light field image is compressed, the compressed light field image may be distorted, and a rough score corresponding to each light field image may be determined according to the degree of distortion of each light field image.
[0077] To distinguish sample light field images with different distortion degrees and ensure that light field images with different distortion degrees have corresponding sample distortion scores. In the embodiments of the present application, through the Ranking-MOS method, different distortion degree types are set, and each distortion degree type corresponds to a distortion level. For example, if the number of distortion levels is set to 5, and the distortion levels include 1-5, where a distortion level of 1 indicates that the corresponding light field image has the least distortion, and a distortion level of 5 indicates that the corresponding light field image has the most distortion. According to the Ranking-MOS method, light field images with different distortion degrees have corresponding sample distortion scores.
[0078] In the embodiments of the present application, the method for generating sample distortion scores through the Ranking-MOS method is: S p = L cd - L lf + 1, where S p represents the sample distortion score of the corresponding light field image, L cd represents the number of distortion levels, and L lf represents the distortion level of the corresponding light field image; for example, when the distortion level of the light field image is 2, the sample distortion score S p = 5 - 2 + 1 = 4. Another example, when the distortion level of the light field image is 5, the sample distortion score S p = 5 - 5 + 1 = 1. For example, taking compression processing as the image information processing, 5 categories can be set for the compression processing, such as these 5 categories are 0-20%, 20%-40%, 40%-60%, 60%-80%, and 80%-100% respectively. Among them, the light field image corresponding to the 0-20% compression processing has a distortion level of 1, and the light field image corresponding to the 80%-100% compression processing has a distortion level of 5. Exemplarily, if the light field image is subjected to compression processing with a distortion degree corresponding to a compression ratio of 20%, a sample light field image is obtained, and the distortion level corresponding to this sample light field image is defined as level 1. According to the Ranking-MOS method, the sample distortion score corresponding to the sample light field image with a distortion degree of level 1 is determined to be 5 points; another example, if the light field image is subjected to compression processing with a distortion degree corresponding to a compression ratio of 90%, and the distortion degree level of the obtained sample light field image is defined as level 5, then the corresponding sample distortion score is 1 point.
[0079] Through the above method, it is possible to perform image information processing on each light field image with different distortion degrees, and give corresponding sample distortion scores based on the light field images with different distortion degrees, so as to facilitate subsequent use of the light field images with different distortion degrees and sample distortion scores as training data.
[0080] (2) Jointly pre-train a preset model according to the sample light field images and the corresponding sample distortion scores to obtain a pre-trained model.
[0081] The preset model can be a network model for deep learning, and the preset model can include modules of different functional types. For example, it includes a convolutional layer (ConvNet), a feature extraction module (Block), a sub-aperture feature fusion module (SAI-Fusion), a pooling layer module (Pool), a fully connected layer module (FC), a global context perception module (GCP), etc.
[0082] Among them, the convolutional layer (ConvNet) contains learned kernels, which are used to distinguish the features of different images or different information of the same image from each other for extraction. For example, to distinguish different types of feature information in each sub-aperture image contained in the light field image.
[0083] The feature extraction module (Block) is used to extract the features of different images or the image information of the same image. For example, in this embodiment, it extracts the sub-aperture features in each sub-aperture image contained in the light field image.
[0084] The sub-aperture feature fusion module (SAI-Fusion) is used to fuse the features. For example, in this embodiment, it fuses the sub-aperture features of the extracted light field image.
[0085] The pooling layer module (Pool) is used to perform downsampling on the fused features. The downsampling process can include max downsampling, average downsampling, etc. For example, in this embodiment, it performs average downsampling on the sub-aperture fusion features obtained by fusing the sub-aperture features of the light field image, and converts the multi-dimensional array matrix corresponding to each sub-aperture fusion feature into a one-dimensional array matrix.
[0086] The fully connected layer module (FC) contains multiple layers of nodes. Among them, each node is connected to all the nodes in the previous layer, and is used to synthesize the previously extracted features to obtain the corresponding result. For example, it synthesizes all the sub-aperture fusion features under the same light field image to obtain the distortion parameter of the light field image.
[0087] The global context perception module (GCP) is used to determine the numerical result according to the association between global features. For example, in this embodiment, it determines the association parameter according to the association or degree of association between all the sub-aperture fusion features under the same light field image.
[0088] In order to improve the accuracy of the target model in recognizing light field images subsequently, first, in this embodiment, a large number of light field images of different information types are obtained, and the image information of each light field image is processed, such as performing compression processing on each light field image to different degrees to obtain sample light field images with different distortion degrees, and performing corresponding sample distortion scoring on the sample light field images with different distortion degrees according to the Ranking-MOS method; further, the sample light field images and the corresponding sample distortion scores are used to pre-train a preset model to obtain a pre-trained model. In this way, it is convenient to further optimize the pre-trained model subsequently to obtain the subsequent target model, so as to improve the accuracy of the target model in recognizing light field images.
[0089] (3) Obtain a real dataset, where the real dataset includes light field images with different compression ratios and the real distortion score corresponding to each light field image.
[0090] Among them, the real dataset includes light field images with different compression ratios and the real distortion score corresponding to each light field image. The real distortion score can be a score of the distortion degree of the light field image according to the standard of human perception.
[0091] In order to obtain a more accurate real dataset, the light field image can be compressed according to an accurate compression ratio to obtain a light field image with an accurate distortion degree, and the range value of the distortion score corresponding to the light field image with this distortion degree is set, and the real distortion score is obtained by scoring the distorted light field image according to human perception. For example, when the light field image is compressed according to a compression ratio of 22%, the real distortion score corresponding to its distortion degree can be 4.9 points. Since there may be small differences in human perception of different light field images and different distortion degrees, the real distortion score can also be set to any value within the score interval value (4.8, 5.0). Another example is when the light field image is compressed according to a compression ratio of 66%, the real distortion score corresponding to its distortion degree can be 2.3 points, and the real distortion score can be any value within the score interval value (2.1, 2.4).
[0092] It should be noted that when obtaining the real dataset in the embodiment of the present application, the same or different light field images can also be compressed according to a 1% compression ratio difference, such as obtaining light field images with a compression ratio of 1% - 100%, and obtaining the real distortion scores of the light field images with the corresponding compression ratios. For example, the real score of the light field image with a distortion degree of 1% is 2.05 points, and another example is that the real score of the light field image with a distortion degree of 2% is 2.10 points, and the real score of the light field image with a distortion degree of 19% is 2.95 points, etc. In this way, an accurate real dataset is obtained, which is convenient for subsequent training of the pre-trained model, realizing fine-tuning of the network parameters of the pre-trained model, and improving the accuracy of the target model in recognizing light field images.
[0093] (4) Jointly train the pre-trained model based on light field images with different compression ratios and their corresponding true distortion scores to obtain the target model.
[0094] After obtaining the pre-trained model in the embodiments of the present application, in order to obtain a more accurate target model, it is necessary to jointly train the pre-trained model with light field images with different precise compression ratios and their true distortion scores, and iteratively train until the model converges to obtain the target model.
[0095] Among them, the step of "jointly training the pre-trained model based on light field images with different compression ratios and their corresponding true distortion scores to obtain the target model" includes:
[0096] (4.1) Input light field images with different compression ratios into the pre-trained model to obtain corresponding predicted distortion scores;
[0097] (4.2) Obtain the difference value between the predicted distortion score and the true distortion score;
[0098] (4.3) Iteratively train the network parameters of the pre-trained model based on the difference value until the difference value converges to obtain the trained target model.
[0099] In order to improve the accuracy of the model in identifying light field images, by inputting light field images with different precise compression ratios into the pre-trained model, the model outputs corresponding predicted distortion scores; by comparing the predicted distortion scores with the true distortion scores, and when there is a difference between the predicted distortion scores and the true distortion scores, obtain the difference value between the two to adjust the network parameters in the pre-trained model according to the difference value; further, continuously input light field images with different precise compression ratios into the pre-trained model, and adjust the network parameters of the pre-trained model according to the difference value between the obtained predicted distortion scores and the true distortion scores until the model converges to obtain the target model.
[0100] In the above way, by using the real dataset to train the pre-trained model, the target model is obtained to facilitate improving the accuracy in subsequent identification of light field images.
[0101] (5) Then the step of "receiving the light field image to be recognized" includes: receiving the light field image to be recognized through the target model.
[0102] In order to improve the efficiency of distortion recognition of the light field image to be recognized, in the embodiments of the present application, after training the target model, the light field image to be recognized can be input into the target model, so that the corresponding module in the target model receives the light field image to facilitate further image recognition processing subsequently to obtain the target distortion score of the light field image to be recognized.
[0103] 102. Extract features from each light field sub-image to obtain a sub-aperture feature set corresponding to each light field sub-image.
[0104] Among them, the light field sub-image can be a natural image of the corresponding object recorded at the corresponding light field angle, and each light field sub-image contains multiple features. For example, a light field sub-image contains multiple objects or image information (such as a car, a person, a road, etc.), each object or image information corresponds to an object type, and the image information of each object type is composed of multiple pixels, and each pixel corresponds to a sub-aperture feature.
[0105] Among them, the sub-aperture feature set contains sub-aperture features of different object types. For example, if each light field sub-image contains image information of object types such as a car, a person, and a road, the sub-aperture feature set contains sub-aperture features of a person, sub-aperture features of a car, and sub-aperture features of a road.
[0106] In order to score the light field image subsequently, each light field sub-image included in the light field image can be subjected to feature extraction to obtain a sub-aperture feature set corresponding to each light field sub-image. Among them, the sub-aperture feature set contains multiple sub-aperture features corresponding to different things.
[0107] In order to improve the recognition efficiency of the light field image, in the embodiment of the present application, when the light field image is recognized by the target model, after the light field image is subjected to convolution operation processing in the convolutional layer of the target model, the light field image can be subjected to feature extraction through the feature extraction module in the target model. In this way, the feature extraction efficiency is improved, and further the recognition efficiency of the light field image is improved.
[0108] 103. Fuse the multiple sub-aperture feature sets according to the object type to obtain sub-aperture fusion features under different object types.
[0109] Among them, the object type can be the type of thing information or image information included in each light field sub-image, and a light field sub-image usually contains image information of multiple object types. For example, each light field sub-image under the same light field image contains image information of multiple object types, and the image information can be a car, a person, a road; among them, the car, the person, and the road respectively correspond to image information of an object type; each image information is composed of multiple pixels, and each pixel corresponds to a sub-aperture feature.
[0110] In order to improve the accuracy when performing distortion recognition on the light field image, in this embodiment, the sub-aperture features of the same object type in each light field sub-image are fused to obtain sub-aperture fusion features under different object types, so as to facilitate subsequent distortion recognition according to the sub-aperture fusion features under different object types, and realize distortion recognition by combining multiple light field sub-images included in the light field image, and improve the accuracy of distortion recognition.
[0111] In some embodiments, the step of "fusing the multiple sub-aperture feature sets according to the object type to obtain sub-aperture fusion features under different object types" includes:
[0112] (1) Identifying the object type of each sub-aperture feature in each sub-aperture feature set.
[0113] Wherein, the sub-aperture feature set contains multiple sub-aperture features corresponding to different object types. For example, the light field sub-image contains image information of three object types: people, vehicles, and roads. The image information of each object type contains multiple pixels, and each pixel corresponds to a sub-aperture feature.
[0114] To improve the efficiency of subsequent feature fusion, in this embodiment, the object type of each sub-aperture feature in each sub-aperture feature set is identified. For example, the light field sub-image contains image information of three object types: people, vehicles, and roads. Each image information contains multiple sub-aperture features of the corresponding object type. When identifying the object type of each sub-aperture feature, the coordinate information of each sub-aperture feature (or pixel) in the original corresponding light field sub-image can be obtained, and the object type of the sub-aperture feature can be determined according to this coordinate information, so as to determine the object type of the sub-aperture feature. Through the above method, the object type of each sub-aperture feature in each sub-aperture feature set can be identified, which is convenient for improving the efficiency of subsequent feature processing.
[0115] (2) Obtaining the color channels of the sub-aperture features corresponding to each object type.
[0116] The color channels can be the color channels included in the sub-aperture features in each light field sub-image, such as the red, green, and blue color channels, or the component colors of red, green, and blue as color channels, which are not limited here. Exemplarily, taking the red, green, and blue color channels as an example, each light field sub-image contains sub-aperture features of multiple object types. Each sub-aperture feature represents the information features in the image information of the corresponding object type. Among them, each sub-aperture feature contains three sub-aperture sub-features corresponding to the red, green, and blue color channels.
[0117] For another example, taking the component colors of red, green, and blue as color channels as an example, each light field sub-image contains sub-aperture features of multiple object types. Each sub-aperture feature represents the information features in the image information of the corresponding object type. The color channel of each sub-aperture feature is the component color of red, green, and blue.
[0118] To improve the accuracy of distortion recognition of light field images, it is necessary to ensure the color channel compatibility during the fusion between sub-aperture features, so as to further determine whether the light field image is distorted. Among them, when obtaining the color channels of the sub-aperture features corresponding to each object type, the obtaining method can be: obtaining the pixel values corresponding to each sub-aperture feature, and determining the color channels of each sub-aperture feature according to the pixel values. In this way, the accuracy during the subsequent fusion of sub-aperture features is ensured.
[0119] (3)Fuse the sub-aperture features of the same object type between the sub-aperture feature sets according to the screening rules of color channels to obtain the sub-aperture fusion features under different object types.
[0120] Among them, the screening rules of the color channels include the same color channel screening and different color channel screening.
[0121] Among them, the sub-aperture fusion features include sub-aperture features of different feature dimensions under the same object type.
[0122] To fuse the sub-aperture features between light field sub-images, in this embodiment, the sub-aperture features of the same object type between different light field sub-images are fused. It should be noted that each sub-aperture feature of the same object type has a corresponding feature dimension, and this feature dimension can be the position information of the corresponding sub-aperture feature under the image information of this object type. For example, taking a car as the object type, the light field sub-image contains multiple sub-aperture features corresponding to the car. For example, multiple adjacent sub-aperture features can form the "rearview mirror", "tire", "body", "windshield", etc. of the car. Among them, each sub-aperture feature has a corresponding feature dimension, such as the feature dimension of this sub-aperture feature can be a certain pixel position in the tire of the car, such as a certain pixel position of the screw or tread in the tire of the car. During the feature fusion in this embodiment, by fusing the sub-aperture features of the same feature dimension under the same object type between each sub-aperture feature set, the sub-aperture fusion features of each object type are obtained, such as the sub-aperture fusion features of the car.
[0123] In some embodiments, step (3) "Fuse the sub-aperture features of the same object type between the sub-aperture feature sets according to the screening rules of color channels to obtain the sub-aperture fusion features under different object types" includes:
[0124] (3.1)Fuse the sub-aperture features with the same color channel between the sub-aperture feature sets to obtain the first sub-aperture fusion features corresponding to different color channels.
[0125] Among them, the first sub-aperture fusion feature is a fusion feature including the same feature dimension and the same color channel. The fusion feature can be a single feature, or can also be a feature set or a feature cluster. It should be noted that each sub-aperture feature of the same object type has a corresponding feature dimension, and the feature dimension can be the position information of the corresponding sub-aperture feature under the image information of the object type.
[0126] (3.2) Filter the first sub-aperture fusion feature to obtain a filtered second sub-aperture fusion feature.
[0127] Since each light field sub-image records the information of the same object at different light field angles, this may cause each light field sub-image to contain sub-aperture features that other light field sub-images do not have. For example, the object types or feature dimensions between the sub-aperture features at the edge parts of each light field sub-image are different. Therefore, after fusing the sub-aperture features with the same color channel under the same feature dimension between the sub-aperture feature sets, there may be some features that cannot be fused, or the number of features participating in the fusion in the obtained first sub-aperture fusion feature is small, indicating that the fusion degree of the first sub-aperture fusion feature is low.
[0128] In order to improve the accuracy of the light field image during distortion recognition in the embodiments of the present application, it is necessary to filter the first sub-aperture fusion feature with a low fusion degree to obtain a filtered second sub-aperture fusion feature. The second sub-aperture fusion feature is obtained by fusing a certain number of sub-aperture features and can have high reference value during the distortion recognition of the light field image. Among them, the method of filtering the first sub-aperture fusion feature can be: obtaining the number of sub-aperture features included in each first sub-aperture fusion feature; filtering the first sub-aperture fusion features with the number of sub-aperture features less than the preset fusion feature number threshold to obtain a filtered second sub-aperture fusion feature.
[0129] (3.3) Fuse the second sub-aperture fusion features of different color channels to obtain sub-aperture fusion features under different object types.
[0130] In order to improve the efficiency of subsequent distortion recognition of the light field image in the embodiments of the present application, after obtaining the second sub-aperture fusion feature, it is necessary to fuse the second sub-aperture fusion features of different color channels. Specifically, the method of fusing the second sub-aperture fusion features of different color channels is: fusing the second sub-aperture fusion features with the same feature dimension under different color channels to obtain sub-aperture fusion features under different object types, where the sub-aperture fusion feature includes sub-aperture features with different feature dimensions under the same object type. In this way, the sub-aperture fusion features under different object types can be used for subsequent distortion recognition of the light field image, and the efficiency of the light field image during distortion recognition can be improved.
[0131] For example, each sub - light - field image in the light - field image contains image information of three object types: people, vehicles, and roads. Each image information contains multiple pixels, each pixel corresponds to a sub - aperture feature, and each sub - aperture feature has a corresponding feature dimension. The feature dimension represents the sub - aperture feature of a specific part in the image information. For example, for the sub - aperture feature corresponding to pixel A in a specific part of a car door, its corresponding feature dimension is this pixel A. Feature fusion is performed on sub - aperture features with the same object type, the same feature dimension, and the same color channel among the sets of sub - aperture features of different sub - light - field images, and filtering is performed after fusion to obtain the second sub - aperture fusion feature. Feature fusion is performed on the second sub - aperture fusion features with the same object type, the same feature dimension, and different color channels to obtain the sub - aperture fusion feature with the same feature dimension under each object type.
[0132] Through the above method, sub - aperture fusion features under different object types are obtained, which is convenient for subsequent use in light - field image distortion recognition and improves the accuracy in subsequent light - field image distortion recognition.
[0133] 104. Generate corresponding correlation parameters according to the correlation degree between sub - aperture fusion features under different object types, and generate distortion parameters according to the recognition degree of sub - aperture fusion features under different object types.
[0134] Among them, the correlation degree represents the correlation degree between any two sub - aperture fusion features, that is, the correlation between any two sub - aperture fusion features. That is, the correlation degree is generated from the correlation between sub - aperture fusion features of different object types. For example, the image information of the object types in each sub - light - field image of the light - field image mainly includes image information of people, vehicles, and roads. By obtaining the correlation degree between sub - aperture fusion features of different feature dimensions of vehicles, the correlation degree between sub - aperture fusion features of different feature dimensions of people, the correlation degree between sub - aperture fusion features of different feature dimensions of roads, and obtaining the correlation degree between the sub - aperture fusion feature of a vehicle and the sub - aperture fusion feature of a person, the correlation degree between the sub - aperture fusion feature of a vehicle and the sub - aperture fusion feature of a road, and the correlation degree between the sub - aperture fusion feature of a person and the sub - aperture fusion feature of a road. Specifically, the process of obtaining the correlation degree is as follows: identify the distance value between sub - aperture fusion features under different object types; determine the correlation coefficient between different sub - aperture fusion features according to the distance value, where the correlation coefficient represents the correlation between different sub - aperture fusion features; obtain the average correlation coefficient corresponding to all correlation coefficients; look up the correlation degree corresponding to this average correlation coefficient from a preset correlation coefficient table, where the preset correlation coefficient table contains the mapping relationship between the correlation coefficient and the correlation degree.
[0135] In order to obtain the correlation between sub-aperture fusion features, this embodiment generates correlation parameters between sub-aperture fusion features of different object types according to the correlation between any two sub-aperture fusion features or multiple sub-aperture fusion features under different object types, so as to facilitate the subsequent participation in the distortion degree and distortion score of the light field image according to the correlation parameters.
[0136] The recognition degree is generated by the difference between sub-aperture fusion features under different object types, or the recognition degree can be the difference between sub-aperture fusion features under different object types and sub-aperture features of the same feature dimension, and the recognition degree reflects the recognizability of sub-aperture fusion features of different object types when generating corresponding images. Specifically, the generation process of the recognition degree can be: based on sub-aperture fusion features under different object types, a feature difference value between adjacent sub-aperture fusion features is obtained, wherein the feature difference value is generated by the color channel difference and the object type difference between any two sub-aperture fusion features, and the feature difference value represents the color channel difference degree and the object type difference degree between any two sub-aperture fusion features; and the recognition degree between sub-aperture fusion features under different object types is determined according to the feature difference value. It should be noted that the recognition degree is generated based on the degree of difference between the sub-aperture fusion features under different object types. The degree of difference can represent the degree of difference in color channels and the degree of difference in object types between two sub-aperture fusion features, such as the degree of difference in color channels and the degree of difference in object types between two adjacent sub-aperture fusion features. Through the degree of difference in color channels and the degree of difference in object types between two adjacent sub-aperture fusion features, the smoothness between the corresponding pixels of the two sub-aperture fusion features when constituting the image can be obtained, and the smoothness can reflect whether the image has distortion.
[0137] In order to obtain the distortion coefficient of the light field image, this embodiment generates distortion parameters of sub-aperture fusion features of different object types according to the recognition degree of sub-aperture fusion features under different object types, so as to subsequently participate in the distortion degree and distortion score of the light field image according to the distortion parameters.
[0138] In some embodiments, distortion recognition can be performed on a light field image using a trained target model, wherein the target model includes a global context perception module (GCP), and the sub-aperture fusion features under different object types are combined by the global context perception module to generate corresponding association parameters. The step of "generating corresponding association parameters according to the degree of association between sub-aperture fusion features under different object types" includes:
[0139] (1) Input the sub-aperture fusion features of different object types into the global context perception module in the trained target model;
[0140] (2) Output the correlation parameter through the global context awareness module. The correlation parameter is generated by the global context awareness module according to the correlation degree between the sub-aperture fusion features of different object types.
[0141] In order to obtain the correlation parameter between the sub-aperture fusion features of different object types, in the embodiments of the present application, the sub-aperture fusion features of different object types are input into the global context awareness module in the trained target model together. The output calculation formula of the global context awareness module is as follows:
[0142]
[0143] Among them, let X represent the sub-aperture fusion features of different object types input into the GCP module. The shape of the sub-aperture fusion features of different object types is N*S, where N represents the number of sub-aperture fusion features, and S represents the color channels of the sub-aperture fusion features. F gcp represents the fully connected layer, and Y represents the correlation parameter output by the GCP module. It should be noted that the size of the fully connected layer is equal to the number of sub-aperture fusion features.
[0144] In some embodiments, before the step of "generating the corresponding correlation parameter according to the correlation degree between the sub-aperture fusion features under different object types", it includes:
[0145] A. Obtain the number of sub-aperture features included in the sub-aperture fusion feature corresponding to each object type.
[0146] Among them, the sub-aperture fusion feature is a fusion feature obtained by sequentially fusing the sub-aperture features with the same color channel and different color channels under the same object type.
[0147] Among them, the number of sub-aperture features is the number of sub-aperture features included in the sub-aperture fusion feature or the number of sub-aperture features participating in the fusion. The number of sub-aperture features can reflect the reliability when identifying the distortion of the light field image according to the sub-aperture fusion feature. It can be understood that when the number of sub-aperture features included in the sub-aperture fusion feature is small, it means that the reference of the sub-aperture fusion feature in participating in the distortion identification of the light field image is low. When the number of sub-aperture features included in the sub-aperture fusion feature is large, it means that the sub-aperture fusion feature has a high reference or a certain reference in participating in the distortion identification of the light field image.
[0148] B. Filter the sub-aperture fusion features with the number of sub-aperture features less than the preset feature number threshold to obtain the filtered sub-aperture fusion features.
[0149] To improve the accuracy of subsequent distortion recognition of light field images, it is necessary to obtain sub-aperture fusion features with high reference. In the implementation of this application, sub-aperture fusion features with a low number of sub-aperture features are filtered. For example, sub-aperture fusion features with a number of sub-aperture features less than a preset feature number threshold are filtered to obtain filtered sub-aperture fusion features.
[0150] C. Perform downsampling on the filtered sub-aperture fusion features to obtain target sub-aperture fusion features.
[0151] To obtain sub-aperture fusion features that meet certain specifications, the embodiments of this application need to perform downsampling on the filtered sub-aperture fusion features. Among them, the downsampling process is to sample some of the features in the filtered sub-aperture fusion features. For example, sample multiple adjacent sub-control fusion features for maximum sampling, or perform average sampling, etc., and fuse the sampled features to obtain target sub-aperture fusion features.
[0152] D. Then generate corresponding correlation parameters according to the correlation degree between sub-aperture fusion features under different object types, and generate distortion parameters according to the recognition degree of sub-aperture fusion features under different object types, including: generating corresponding correlation parameters according to the correlation degree between target sub-aperture fusion features of different object types, and generating distortion parameters according to the recognition degree of target sub-aperture fusion features of different object types.
[0153] In this embodiment, the sub-aperture fusion features of different object types can be input together into the feature extraction module in the trained target model, so that the feature extraction module filters the sub-aperture fusion features with a number of sub-aperture features less than the preset feature number threshold to obtain filtered sub-aperture fusion features; further, input the filtered sub-aperture fusion features into the pooling module or pooling layer in the target model, and perform pooling processing (downsampling processing) through the pooling module or pooling layer. This pooling processing can be average pooling processing to obtain target sub-aperture fusion features. Through the above method, corresponding target sub-aperture fusion features can be obtained, so as to subsequently determine correlation parameters and distortion parameters according to the target sub-aperture fusion features and improve the accuracy of distortion recognition of light field images.
[0154] 105. Determine the target distortion score of the light field image according to the correlation parameter and the distortion parameter.
[0155] Among them, the target distortion score can numerically reflect the distortion degree of the light field image and can convey the distortion situation of the light field image to the user.
[0156] To obtain a relatively accurate target distortion score for the light field image, in the embodiments of the present application, the target distortion score of the light field image is determined by using the correlation parameter and the distortion parameter obtained in the embodiments of the present application. Specifically, the correlation parameter and the distortion parameter can be multiplied to obtain the target distortion score. In this way, the distortion parameter is limited according to the correlation between the sub-aperture fusion features under different object types, and thus the determined target distortion score is obtained. It can be understood that when the correlation between the sub-aperture fusion features under different object types is high, the correlation parameter is large, and the obtained target distortion score is high; when the correlation between the sub-aperture fusion features under different object types is low, the correlation parameter is small, and the obtained target distortion score is small.
[0157] As can be seen from the above, the embodiments of the present application can receive a light field image to be recognized, where the light field image includes light sub-images at different light field angles; extract features from each light sub-image to obtain a set of sub-aperture features corresponding to each light sub-image; fuse the multiple sets of sub-aperture features according to the object type to obtain sub-aperture fusion features under different object types; generate a corresponding correlation parameter according to the correlation degree between the sub-aperture fusion features under different object types, and generate a distortion parameter according to the recognition degree of the sub-aperture fusion features under different object types, where the correlation degree is generated by the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated by the difference degree between the sub-aperture fusion features under different object types; determine the target distortion score of the light field image according to the correlation parameter and the distortion parameter. Thus, by fusing the sub-aperture features of the same object type among the light sub-images in the light field image to obtain sub-aperture fusion features, determining the correlation coefficient according to the correlation features among the sub-aperture fusion features, and determining the distortion coefficient according to the sub-aperture fusion features, and further determining the target quality score of the light field image. In this way, the distortion recognition is realized by combining the global sub-aperture features in the light field image, and the light field image with any distortion degree can be accurately recognized, and the distortion degree of the light field image can be accurately determined, the distortion degree of the image can be accurately recognized, and the accuracy of image recognition can be improved, which has reliability.
[0158] As can be seen from the above, when the embodiment of the present application identifies the distortion of the light field image, by fusing the sub-aperture features of the same feature dimension under the same corresponding type between the light field sub-images, the sub-aperture fusion features of different feature dimensions under different object types are obtained, the corresponding correlation coefficients are generated according to the correlation between the sub-aperture fusion features of different feature dimensions under different object types, and the corresponding distortion coefficients are generated through the fully connected layer according to the sub-aperture fusion features of different feature dimensions under different object types. Furthermore, the target distortion score of the light field image is determined according to the correlation coefficient and the distortion coefficient, so that the user can learn about the distortion situation of the current light field image. In the above way, the distortion of the light field image is identified by combining the global sub-aperture fusion features, the distortion degree of the light field image is accurately determined, the distortion degree of the image is accurately identified, and the accuracy of image recognition is improved.
[0159] According to the method described in the above embodiment, the following will be further described in detail by way of examples.
[0160] Taking the distortion recognition of the light field image as an example, the image recognition method provided by the embodiment of the present application will be further described.
[0161] See Figure 3 , Figure 3 which is another step flow diagram of the image recognition method provided by the embodiment of the present application, Figure 4 which is a structural diagram of the preset model provided by the embodiment of the present application, Figure 5 which is a schematic diagram of the sub-aperture feature fusion process provided by the embodiment of the present application; for ease of understanding, please also combine Figure 3 , Figure 4 and Figure 5 to describe the embodiment of the present application.
[0162] In the embodiment of the present application, it will be described from the perspective of an image recognition device, which can be specifically integrated in computer devices such as terminals, servers and other devices. When the processors on the terminal and the server jointly execute the program corresponding to the image recognition method, the specific timing process corresponding to the image recognition method is as follows:
[0163] 201. Obtain a sample light field image set, which includes a plurality of distorted sample light field images and the corresponding sample distortion scores for each sample light field image.
[0164] Among them, the light field image can be an image describing the light distribution information in space. Through the light field image, the intuitive information of the same thing in space at different positions and different angles can be recorded, that is, the light field image includes a plurality of light field sub-images with different light field angles, and each light field sub-image can be a sub-image recording the corresponding position and light angle.
[0165] In order to distinguish sample light field images with different distortion degrees and ensure that light field images with different distortion degrees have corresponding sample distortion scores. In the embodiments of the present application, through the ranking label method (Ranking-MOS), different distortion degree types are set, and each distortion degree type corresponds to a distortion level. For example, if the number of distortion levels is set to 5, and the distortion levels include 1-5, where a distortion level of 1 indicates that the corresponding light field image has the least distortion, and a distortion level of 5 indicates that the corresponding light field image has the most distortion. According to the ranking label method (Ranking-MOS), light field images with different distortion degrees are generated with corresponding sample distortion scores.
[0166] For example, the method for generating sample distortion scores through the Ranking-MOS method is: S p = L cd - L lf + 1, where S p represents the sample distortion score of the corresponding light field image, L cd represents the number of distortion levels, and L lf represents the distortion level of the corresponding light field image; for example, when the distortion level of the light field image is 2, the sample distortion score S p = 5 - 2 + 1 = 4. Another example, when the distortion level of the light field image is 5, the sample distortion score S p = 5 - 5 + 1 = 1. For example, taking compression processing as the image information processing, the compression processing can be set to 5 categories, such as these 5 categories are 0-20%, 20%-40%, 40%-60%, 60%-80%, and 80%-100% respectively. Among them, the light field image corresponding to the 0-20% compression processing has a distortion level of 1, and the light field image corresponding to the 80%-100% compression processing has a distortion level of 5. Exemplarily, if the light field image is subjected to compression processing with a distortion degree corresponding to a compression ratio of 20%, a sample light field image is obtained, and the distortion level corresponding to this sample light field image is defined as level 1. According to the Ranking-MOS method, the sample distortion score corresponding to the sample light field image with a distortion degree of level 1 is determined to be 5 points; another example, if the light field image is subjected to compression processing with a distortion degree corresponding to a compression ratio of 90%, and the distortion degree level of the obtained sample light field image is defined as level 5, then the corresponding sample distortion score is 1 point.
[0167] Through the above method, it is possible to perform image information processing on each light field image with different distortion degrees and give corresponding sample distortion scores based on the light field images with different distortion degrees, so as to facilitate subsequent use of the light field images with different distortion degrees and sample distortion scores as training data.
[0168] 202. Jointly pre-train a preset model according to the sample light field image and the corresponding sample distortion score to obtain a pre-trained model.
[0169] Among them, the preset model is a basic model for distortion recognition of light field images. As Figure 4 shown, the preset model 40 includes: a first convolutional layer (ConvNet) 41, a first feature extraction module (Block1-7) 42, a sub-aperture feature fusion module (SAI-Fusion) 43, a second feature extraction module (Block9-13) 44, a pooling layer module (Pool) 45, a second convolutional layer (ConvNet) 46, a fully connected layer module (FC) 47, a global context perception module (GCP) 48, and an operation module 49.
[0170] Among them, the convolutional layer (ConvNet) contains learned kernels that are used to distinguish the features of different images or different information of the same image from each other for extraction. For example, the different types of feature information in each light field sub-image included in the light field image are distinguished.
[0171] The feature extraction module (Block) is used to extract the features of different images or the image information of the same image. In this embodiment, it is used to extract the sub-aperture features in each light field sub-image included in the light field image.
[0172] The sub-aperture feature fusion module (SAI-Fusion) is used to fuse the features. Among them, the sub-aperture feature fusion module includes a first sub-aperture feature fusion block (SAI-Fusion1), a sub-aperture feature extraction block (Block8), and a second sub-aperture feature fusion block (SAI-Fusion2). In this embodiment, the sub-aperture features of the extracted light field image are fused. See Figure 5 . It is a schematic diagram of the sub-aperture feature fusion module fusing sub-aperture features. In order to obtain the relationship between the sub-aperture features (SAI) of the light field image, the embodiment of the present application fuses the features of SAI through the SAI-Fusion module. Among them, the sub-aperture features under different object types are set as F ij , where i represents the number of SAI features, and j represents the number of channels of each feature. Assume that the number of SAI features is N, and the number of channels of each feature is M, that is, i ∈ {1, 2, 3, 4, 5,..., n} and j ∈ {1, 2, 3, 4, 5,..., m}. The SAI-Fusion module can be expressed as: F' ij = F sai-fusion (F 1i , F 2i , … F ni ) j , where F sai-fusionRepresents a series of convolutional operations that complete the feature fusion of the Fsai-fusion module with relatively low computational complexity by adopting the MobilenetV3 Block. In SAI-Fusion1, the same channels of each SAI feature are fused. In SAI-Fusion2, different channels of the SAI features are fused together. It should be noted that the number of convolutional kernels in the last convolutional layer of SAI-Fusion1 is equal to the number N of SAI. Assuming that the size of the sub-aperture features input to the sub-aperture feature fusion module (SAI-Fusion) is N×M×H×W, where H and W represent the height and width of the input features respectively, the size of the features output by SAI-Fusion1 is M×N×H×W. Since the sub-aperture feature extraction block (8) does not change the size of the features, the size of the input features of SAI-Fusion2 is M×N×H×W, and the number of convolutional kernels in the last convolutional layer of SAI-Fusion2 is equal to the input channel M of SAI-Fusion1, so the size of its output features is N×M×H×W.
[0173] The pooling layer module (Pool) is used to perform downsampling on the fused features. This downsampling process can include max-pooling, average-pooling, etc.; in this embodiment, for example, average downsampling is performed on the sub-aperture fused features obtained by fusing the sub-aperture features of the light field image, and the multi-dimensional array matrix corresponding to each sub-aperture fused feature is converted into a one-dimensional array matrix.
[0174] The fully connected layer module (FC) contains multiple layers of nodes. Among them, each node is connected to all nodes in the previous layer and is used to synthesize the previously extracted features to obtain the corresponding results. For example, all the sub-aperture fused features under the same light field image are synthesized to obtain the distortion parameters of the light field image.
[0175] The global context perception module (GCP) is used to determine the numerical results based on the associations between global features. For example, in this embodiment, the association parameters are determined according to the associations or degrees of association between all the sub-aperture fused features under the same light field image.
[0176] In order to improve the accuracy of the target model in recognizing light field images in the following, first, in this embodiment, a large number of light field images of different information types are obtained, and each light field image is subjected to image information processing, such as performing compression processing on each light field image to different degrees to obtain sample light field images with different distortion degrees, and performing corresponding sample distortion scoring on the sample light field images with different distortion degrees according to the Ranking-MOS method; further, the sample light field images and the corresponding sample distortion scores are used to pre-train a preset model to obtain a pre-trained model.
[0177] 203. Obtain a real dataset, where the real dataset includes light field images with different compression ratios and the corresponding real distortion scores for each light field image.
[0178] Among them, the real dataset includes light field images with different compression ratios and the corresponding real distortion scores for each light field image. The real distortion score can be a score for the distortion degree of the light field image according to the criteria of human perception.
[0179] In order to obtain a more accurate real dataset, the light field images can be compressed according to precise compression ratios to obtain light field images with precise distortion degrees, and the range of distortion scores corresponding to the light field images with such distortion degrees can be set, and the distorted light field images can be scored according to human perception to obtain real distortion scores. For example, in the embodiment of the present application, when obtaining the real dataset, the same or different light field images can be compressed according to a compression ratio difference of 1%, such as obtaining light field images with compression ratios from 1% to 100%, and obtaining the real distortion scores of the light field images with the corresponding compression ratios. For example, the real score of the light field image with a distortion degree of 1% is 2.05 points, and another example is that the real score of the light field image with a distortion degree of 2% is 2.10 points, and the real score of the light field image with a distortion degree of 19% is 2.95 points, etc. In this way, an accurate real dataset can be obtained for subsequent training of the pre-trained model to improve the accuracy of the target model in recognizing light field images.
[0180] 204. Jointly train the pre-trained model according to the light field images with different compression ratios and the corresponding real distortion scores to obtain a trained target model.
[0181] In order to obtain a more accurate target model, it is necessary to jointly train the pre-trained model with light field images with different precise compression ratios and real distortion scores, and iterate the training until the model converges to obtain the target model.
[0182] Specifically, when jointly training the pre-trained model according to the light field images with different compression ratios and the corresponding real distortion scores, the training process can be as follows: input the light field images with different compression ratios into the pre-trained model to obtain the corresponding predicted distortion scores, obtain the difference value between the predicted distortion scores and the real distortion scores, fine-tune the network parameters of the pre-trained model based on the difference value, and perform iterative training until the difference value converges to obtain the trained target model. It can be understood that the structure of the neural network of the target model is the same as that of the preset model, and the present embodiment will not be shown in the drawings.
[0183] In the above manner, by using the real dataset to train the pre-trained model, a trained target model is obtained to facilitate improving the accuracy in subsequent recognition of light field images.
[0184] 205. Receive the light field image to be recognized through the convolutional layer of the target model, and perform convolutional processing on the light field sub-images included in the light field image with different light field angles.
[0185] Among them, the light field image includes light field sub-images with different light field angles. This convolutional processing can be to extract each light field sub-image in the light field image so that each light field is separated from the light field image.
[0186] In addition, this convolutional processing can also perform convolutional processing on the features included in each light field sub-image in the light field image to obtain the features corresponding to each light field sub-image. For example, by identifying the pixels in each light field sub-image, feature extraction is performed on each pixel to obtain the features corresponding to each light field sub-image.
[0187] 206. Perform feature extraction on each light field sub-image through the first feature extraction module of the target model to obtain a sub-aperture feature set corresponding to each light field sub-image.
[0188] The first feature extraction module includes multiple feature extraction blocks (Blocks). The embodiments of this application include 7 feature extraction blocks, such as Block1, Block2, Block3, Block4, Block5, Block6, and Block7. Through the cooperation of the above feature extraction blocks, sub-aperture feature extraction is performed to obtain a sub-aperture feature set corresponding to each light field sub-image.
[0189] For example, each light field sub-image includes image information of object types such as cars, people, and roads. Then, feature extraction is performed on each light field sub-image included in the light field image to obtain a sub-aperture feature set corresponding to each light field sub-image; among them, the sub-aperture feature set includes sub-aperture features of people, sub-aperture features of cars, and sub-aperture features of roads.
[0190] 207. Through the sub-aperture feature fusion module of the target model, fuse multiple sub-aperture feature sets according to object types to obtain sub-aperture fusion features under different object types.
[0191] Among them, the sub-aperture feature fusion module includes a first sub-aperture feature fusion block (SAI-Fusion1), a sub-aperture feature extraction block (Block8), and a second sub-aperture feature fusion block (SAI-Fusion2).
[0192] To improve the efficiency during subsequent feature fusion, the embodiments of the present application identify the object types of each sub-aperture feature in each sub-aperture feature set. For example, the light field sub-image contains image information of three object types: person, vehicle, and road. Each image information contains multiple sub-aperture features corresponding to the respective object types. When identifying the object type of each sub-aperture feature, the coordinate information of each sub-aperture feature (or pixel) in the original corresponding light field sub-image can be obtained, and based on this coordinate information, the image information to which the sub-aperture feature belongs can be determined, thereby determining the object type of the sub-aperture feature.
[0193] To improve the accuracy of distortion recognition of the light field image, it is necessary to ensure the color channel compatibility during the fusion between sub-aperture features, so as to further determine whether the light field image is distorted. The embodiments of the present application obtain the color channels of the sub-aperture features corresponding to each object type. Among them, the color channel can be the color channel included in the sub-aperture features of each light field sub-image. Taking the component colors of red, green, and blue as an example of the color channel, each light field sub-image contains sub-aperture features of multiple object types. Each sub-aperture feature represents an information feature in the image information of the corresponding object type, and the color channel of each sub-aperture feature is the component color of red, green, and blue. Specifically, obtaining the light field can be: obtaining the pixel value corresponding to each sub-aperture feature, and determining the color channel of each sub-aperture feature according to the pixel value. In this way, the accuracy during subsequent sub-aperture feature fusion is ensured.
[0194] Further, the sub-aperture features of the same object type between sub-aperture feature sets are fused according to the screening rules of the color channels to obtain sub-aperture fusion features under different object types. Among them, the screening rules of the color channels include the same color channel screening and different color channel screening. It should be noted that the sub-aperture fusion feature includes sub-aperture features of different feature dimensions under the same object type.
[0195] In the embodiments of the present application, "fusing the sub-aperture features of the same object type between sub-aperture feature sets according to the screening rules of the color channels to obtain sub-aperture fusion features under different object types" may include:
[0196] (1) Fusing the sub-aperture features with the same color channel between sub-aperture feature sets through the first sub-aperture feature fusion block (SAI-Fusion1) to obtain the first sub-aperture fusion features corresponding to different color channels. Among them, the first sub-aperture fusion feature is a fusion feature including the same feature dimension and the same color channel. The fusion feature can be a single feature, or a feature set or a feature cluster. It should be noted that each sub-aperture feature of the same object type has a corresponding feature dimension, and the feature dimension can be the position information of the corresponding sub-aperture feature in the image information of the object type.
[0197] (2) Filter the first sub-aperture fusion feature through the sub-aperture feature extraction block (Block8) to obtain the filtered second sub-aperture fusion feature.
[0198] Since each light field sub-image records the information of the same object at different light field angles, this may cause each light field sub-image to contain sub-aperture features that other light field sub-images do not have. For example, the object types or feature dimensions between the sub-aperture features at the edge parts of each light field sub-image are different. Therefore, after fusing the sub-aperture features with the same feature dimension and the same color channel between the sub-aperture feature sets, there may be some features that cannot be fused, or the number of features participating in the fusion in the obtained first sub-aperture fusion feature is small, indicating that the fusion degree of the first sub-aperture fusion feature is low.
[0199] In order to improve the accuracy of distortion recognition of light field images in the embodiments of the present application, it is necessary to filter the first sub-aperture fusion feature with a low fusion degree to obtain the filtered second sub-aperture fusion feature. The second sub-aperture fusion feature is obtained by fusing a certain number of sub-aperture features and can have high reference value in the distortion recognition of light field images. Among them, the method of filtering the first sub-aperture fusion feature can be: obtaining the number of sub-aperture features included in each first sub-aperture fusion feature; filtering the first sub-aperture fusion feature with the number of sub-aperture features less than the preset fusion feature number threshold to obtain the filtered second sub-aperture fusion feature. Among them, the preset fusion feature number threshold can be a screening value for the number of sub-aperture features participating in the fusion to obtain the sub-aperture fusion feature, and is used to screen the subsequent sub-aperture fusion features for distortion recognition of light field images. Through the above method, the quality of the subsequent sub-aperture fusion features for distortion recognition of light field images can be ensured, the accuracy of distortion recognition of light field images can be improved, and it has reliability.
[0200] (3) Feature fusion of the second sub-aperture fusion features of different color channels is performed through the second sub-aperture feature fusion block (SAI-Fusion2) to obtain sub-aperture fusion features under different object types. Specifically, the method of fusing the second sub-aperture fusion features of different color channels is: fusing the second sub-aperture fusion features with the same feature dimension under different color channels to obtain sub-aperture fusion features under different object types, where the sub-aperture fusion feature contains sub-aperture features with different feature dimensions under the same object type. In this way, the sub-aperture fusion features under different object types can be used for subsequent distortion recognition of light field images to improve the efficiency of distortion recognition of light field images.
[0201] For example, each light field sub-image in the light field image contains image information of three object types: people, vehicles, and roads. Each image information contains multiple pixels, each pixel corresponds to a sub-aperture feature, and each sub-aperture feature has a corresponding feature dimension. The feature dimension represents the sub-aperture feature of a specific part in the image information. For example, the sub-aperture feature corresponding to pixel A in a specific part of the car door has a corresponding feature dimension that is pixel A. Feature fusion is performed on sub-aperture features with the same object type, the same feature dimension, and the same color channel among the sets of sub-aperture features of different light field sub-images, and filtering is performed after the fusion to obtain a second sub-aperture fusion feature; Feature fusion is performed on the second sub-aperture fusion features with the same object type, the same feature dimension, and different color channels to obtain sub-aperture fusion features with the same feature dimension under each object type.
[0202] 208. Obtain the number of sub-aperture features included in the sub-aperture fusion feature corresponding to each object type through the second feature extraction module of the target model, and filter the sub-aperture fusion features with the number of sub-aperture features less than the preset feature number threshold to obtain the filtered sub-aperture fusion feature.
[0203] Among them, the second feature extraction module includes multiple feature extraction blocks (Block). The embodiment of the present application includes 5 feature extraction blocks, such as Block9, Block10, Block11, Block12, and Block13. Through the cooperation of the above feature extraction blocks, feature extraction (filtering) is performed on the sub-aperture fusion feature corresponding to each object type to obtain the filtered sub-aperture fusion feature.
[0204] Among them, the sub-aperture fusion feature is a fusion feature obtained by sequentially fusing sub-aperture features of the same color channel and different color channels under the same object type.
[0205] Among them, the number of sub-aperture features is the number of sub-aperture features included in the sub-aperture fusion feature or the number of sub-aperture features participating in the fusion. The number of sub-aperture features can reflect the reliability when identifying the distortion of the light field image based on the sub-aperture fusion feature. It can be understood that when the number of sub-aperture features included in the sub-aperture fusion feature is small, it means that the reference of the sub-aperture fusion feature in participating in the distortion identification of the light field image is low. When the number of sub-aperture features included in the sub-aperture fusion feature is large, it means that the sub-aperture fusion feature has a high reference or a certain reference when participating in the distortion identification of the light field image.
[0206] In order to improve the accuracy of subsequent distortion recognition of light field images, it is necessary to obtain sub-aperture fusion features with high reference. In the implementation of this application, sub-aperture fusion features with a low number of sub-aperture features are filtered. For example, sub-aperture fusion features with a number of sub-aperture features less than a preset feature number threshold are filtered to obtain the filtered sub-aperture fusion features.
[0207] 209. Downsample the filtered sub-aperture fusion features through the pooling layer of the target model to obtain the target sub-aperture fusion features.
[0208] In order to obtain sub-aperture fusion features that meet certain specifications, the embodiments of this application need to perform downsampling on the filtered sub-aperture fusion features. Among them, this downsampling process is to sample some features in the filtered sub-aperture fusion features. For example, sample multiple adjacent sub-control fusion features for maximum sampling or average sampling, etc., and fuse the sampled features to obtain the target sub-aperture fusion features. It should be noted that the pooling layer also converts the multi-dimensional array matrix corresponding to the sub-aperture fusion features under different object types into a one-dimensional array matrix.
[0209] For example, input the filtered sub-aperture fusion features into the pooling module or pooling layer in the target model, and perform pooling processing (downsampling processing) through the pooling module or pooling layer. This pooling processing can be average pooling processing. After performing pooling processing on the filtered sub-aperture fusion features, the target sub-aperture fusion features are obtained, so as to input the target sub-aperture fusion features into the global context awareness module and the fully connected layer of the target model at the same time for further processing.
[0210] It should be noted that before the pooling layer performs average pooling processing on the features, the filtered sub-aperture fusion features output by the second feature extraction module can also be convolved through the second convolutional layer in the target model to integrate the filtered sub-aperture fusion features, such as integration in sorting, layout integration, etc., to improve the efficiency of the subsequent pooling layer during pooling processing.
[0211] 210. Generate corresponding correlation parameters through the global context awareness module of the target model according to the correlation degree between the sub-aperture fusion features under different object types; and generate distortion parameters through the fully connected layer of the target model according to the recognition degree of the sub-aperture fusion features under different object types. Among them, the correlation degree is generated by the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated by the difference degree between the sub-aperture fusion features of different object types.
[0212] For example, the sub-aperture fusion features of different object types are input into the global context perception module (GCP) in the trained target model, and the association parameters are output through the global context perception module. The association parameters are generated by the global context perception module according to the association degree between the sub-aperture fusion features of different object types.
[0213] In order to obtain the association parameters between the sub-aperture fusion features of different object types, in the embodiment of the present application, the sub-aperture fusion features of different object types are input together into the global context perception module in the trained target model. The output calculation formula of the global context perception module is as follows:
[0214]
[0215] Among them, let X represent the sub-aperture fusion features of different object types input into the GCP module. The shape of the sub-aperture fusion features of different object types is N*S, where N represents the number of sub-aperture fusion features, and S represents the color channels of the sub-aperture fusion features. F gcp represents the fully connected layer, and Y represents the association parameters output by the GCP module. It should be noted that the size of the fully connected layer is equal to the number of sub-aperture fusion features.
[0216] In addition, the fully connected layer (FC) includes a first fully connected layer (FC1) and a second fully connected layer (FC2). In order to obtain the distortion parameters of the sub-aperture fusion features of different object types, in the embodiment of the present application, they are generated through the fully connected layer (FC) in the target model. For example, the sub-aperture fusion features of different object types are input into the fully connected layer in the trained target model, so that the fully connected layer outputs the corresponding distortion parameters.
[0217] 211. The operation module of the target model determines the target distortion score of the light field image according to the association parameters and the distortion parameters.
[0218] Among them, the operation module multiplies the association parameters and the distortion parameters to obtain the target distortion score of the light field image.
[0219] In order to obtain a more accurate target distortion score of the light field image, in the embodiment of the present application, the target distortion score of the light field image is determined by the association parameters and the distortion parameters obtained in the embodiment of the present application. Specifically, the association parameters and the distortion parameters can be multiplied to obtain the target distortion score. In this way, the distortion parameters are limited according to the association between the sub-aperture fusion features under different object types, so as to determine the target distortion score. It can be understood that when the association between the sub-aperture fusion features under different object types is relatively high, the association parameters are relatively large, and the obtained target distortion score is relatively high; when the association between the sub-aperture fusion features under different object types is relatively low, the association parameters are relatively small, and the obtained target distortion score is relatively small.
[0220] As can be seen from the above, the embodiments of the present application can receive a light field image to be recognized, where the light field image includes light sub-images at different light field angles; perform feature extraction on each light sub-image to obtain a sub-aperture feature set corresponding to each light sub-image; fuse multiple sub-aperture feature sets according to object types to obtain sub-aperture fusion features under different object types; generate corresponding correlation parameters according to the correlation degrees between the sub-aperture fusion features under different object types, and generate distortion parameters according to the recognition degrees of the sub-aperture fusion features under different object types, where the correlation degree is generated by the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated by the difference degree between the sub-aperture fusion features of different object types; determine the target distortion score of the light field image according to the correlation parameters and the distortion parameters. Thus, by fusing the sub-aperture features of the same object type among the light sub-images in the light field image to obtain sub-aperture fusion features, determining the correlation coefficient according to the correlation features between the sub-aperture fusion features, and determining the distortion coefficient according to the sub-aperture fusion features, and further determining the target quality score of the light field image, thereby realizing distortion recognition by combining the global sub-aperture features in the light field image, accurately recognizing light field images with any degree of distortion, accurately determining the degree of distortion of the light field image, precisely identifying the degree of image distortion, improving the accuracy of image recognition, and having reliability.
[0221] To better implement the above method, the embodiments of the present application further provide an image recognition device, which can be specifically integrated in a computer device such as a terminal or a server. The meanings of the nouns are the same as those in the above image recognition method, and the specific implementation details can refer to the description in the method embodiments.
[0222] As Figure 6 shown, Figure 6 is a schematic structural diagram of the image recognition device provided by the embodiments of the present application. The image recognition device may include a receiving unit 301, an extracting unit 302, a fusing unit 303, a generating unit 304, and a determining unit 305, as follows:
[0223] The receiving unit 301 is configured to receive a light field image to be recognized, where the light field image includes light sub-images at different light field angles;
[0224] The extracting unit 302 is configured to perform feature extraction on each light sub-image to obtain a sub-aperture feature set corresponding to each light sub-image;
[0225] The fusing unit 303 is configured to fuse multiple sub-aperture feature sets according to object types to obtain sub-aperture fusion features under different object types;
[0226] A generating unit 304, configured to generate corresponding correlation parameters according to the correlation degrees between the sub-aperture fusion features under different object types, and generate distortion parameters according to the recognition degrees of the sub-aperture fusion features under different object types;
[0227] A determining unit 305, configured to determine a target distortion score of the light field image according to the correlation parameters and the distortion parameters.
[0228] In some embodiments, the fusion unit 303 includes:
[0229] An identifying subunit, configured to identify the object type of each sub-aperture feature in each sub-aperture feature set;
[0230] An obtaining subunit, configured to obtain the color channels of the sub-aperture features corresponding to each object type;
[0231] A fusing subunit, configured to fuse the sub-aperture features of the same object type between the sub-aperture feature sets according to the screening rules of the color channels, so as to obtain sub-aperture fusion features under different object types.
[0232] In some embodiments, the fusing subunit is further configured to:
[0233] Fuse the sub-aperture features with the same color channels between the sub-aperture feature sets to obtain first sub-aperture fusion features corresponding to different color channels;
[0234] Filter the first sub-aperture fusion features to obtain filtered second sub-aperture fusion features;
[0235] Fuse the second sub-aperture fusion features of different color channels to obtain sub-aperture fusion features under different object types.
[0236] In some embodiments, the generating unit is further configured to:
[0237] Fuse the sub-aperture features with the same color channels between the sub-aperture feature sets to obtain first sub-aperture fusion features corresponding to different color channels;
[0238] Filter the first sub-aperture fusion features to obtain filtered second sub-aperture fusion features;
[0239] Fuse the second sub-aperture fusion features of different color channels to obtain sub-aperture fusion features under different object types.
[0240] In some embodiments, the generating unit is further configured to:
[0241] Input the sub-aperture fusion features of different object types into the global context awareness module in the trained target model;
[0242] Output the correlation parameters through the global context awareness module, where the correlation parameters are generated by the global context awareness module according to the correlation degree between the sub-aperture fusion features of different object types.
[0243] In some embodiments, the device further includes: a training unit, configured to:
[0244] Obtain a sample light field image set, where the sample light field image set includes a plurality of distorted sample light field images and the corresponding sample distortion scores for each sample light field image;
[0245] Jointly pre-train a preset model according to the sample light field images and the corresponding sample distortion scores to obtain a pre-trained model;
[0246] Obtain a real dataset, where the real dataset includes light field images with different compression ratios and the corresponding real distortion scores for each light field image;
[0247] Jointly train the pre-trained model according to the light field images with different compression ratios and the corresponding real distortion scores to obtain a target model;
[0248] Then, the receiving unit 301 is further configured to receive the light field image to be recognized through the target model.
[0249] In some embodiments, the training unit is further configured to:
[0250] Input the light field images with different compression ratios into the pre-trained model to obtain the corresponding predicted distortion scores;
[0251] Obtain the difference value between the predicted distortion score and the real distortion score;
[0252] Iteratively train the network parameters of the pre-trained model based on the difference value until the difference value converges to obtain the trained target model.
[0253] In some embodiments, the device further includes: a processing unit, configured to:
[0254] Obtain the number of sub-aperture features included in the sub-aperture fusion feature corresponding to each object type;
[0255] Filter the sub-aperture fusion features with the number of sub-aperture features less than the preset feature number threshold to obtain the filtered sub-aperture fusion features;
[0256] Perform downsampling processing on the filtered sub-aperture fusion features to obtain the target sub-aperture fusion features;
[0257] Then, the generating unit 304 is further configured to generate the corresponding correlation parameters according to the correlation degree between the target sub-aperture fusion features of different object types, and generate distortion parameters according to the recognition degree of the target sub-aperture fusion features of different object types.
[0258] As can be seen from the above, the image recognition device provided by the embodiment of the present application can receive the light field image to be recognized through the receiving unit 301, and the light field image includes sub-aperture images with different light field angles; extract the features of each sub-aperture image through the extraction unit 302 to obtain a sub-aperture feature set corresponding to each sub-aperture image; fuse the multiple sub-aperture feature sets according to the object type through the fusion unit 303 to obtain sub-aperture fusion features under different object types; generate corresponding correlation parameters according to the correlation degree between the sub-aperture fusion features under different object types through the generation unit 304, and generate distortion parameters according to the recognition degree of the sub-aperture fusion features under different object types, wherein the correlation degree is generated by the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated by the difference degree between the sub-aperture fusion features of different object types; determine the target distortion score of the light field image according to the correlation parameters and distortion parameters through the determination unit 305. Thus, by fusing the sub-aperture features of the same object type among the sub-aperture images in the light field image to obtain sub-aperture fusion features, determining the correlation coefficient according to the correlation features between the sub-aperture fusion features, and determining the distortion coefficient according to the sub-aperture fusion features, and further determining the target quality score of the light field image, thereby realizing distortion recognition by combining the global sub-aperture features in the light field image, accurately recognizing the light field image with any degree of distortion, accurately determining the degree of distortion of the light field image, accurately identifying the degree of image distortion, improving the accuracy of image recognition, and having reliability.
[0259] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated herein.
[0260] The embodiment of the present application provides a computer device, specifically referring to Figure 7 , which shows the structural schematic diagram of the computer device involved in the embodiment of the present application. The structure of the computer device is specifically as follows:
[0261] The computer device may include a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, an input unit 404 and other components. Those skilled in the art can understand that Figure 7 the computer device structure shown in
[0262] The processor 401 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and circuits. By running or executing software programs and / or units stored in the memory 402, and by invoking the data stored in the memory 402, it executes various functions of the computer device and processes data, thereby performing an overall detection of the computer device. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 401 either.
[0263] The memory 402 can be used to store software programs and units. The processor 401 executes various functional applications and data processing by running the software programs and units stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as the sound playback function, video playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0264] The computer device further includes a power supply 403 for powering each component. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby implementing functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 403 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0265] The computer device may further include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0266] Although not shown, the computer device may further include a display unit, etc., which will not be elaborated here. Specifically, in this embodiment, the processor 401 in the computer device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to achieve various functions as follows:
[0267] When the computer device is a terminal or a server, the processor 401 may perform the following operations: receiving a light field image to be recognized, where the light field image includes light sub-images at different light field angles; extracting features from each light sub-image to obtain a set of sub-aperture features corresponding to each light sub-image; fusing the multiple sets of sub-aperture features according to the object type to obtain sub-aperture fusion features under different object types; generating corresponding correlation parameters based on the correlation degree between the sub-aperture fusion features under different object types, and generating distortion parameters based on the recognition degree of the sub-aperture fusion features under different object types; and determining the target distortion score of the light field image according to the correlation parameters and the distortion parameters.
[0268] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.
[0269] The present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the image recognition method provided in various optional implementation manners in the above embodiments.
[0270] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. The computer program can be loaded by a processor to execute the steps in any one of the image recognition methods provided in the embodiments of the present application. For example, the computer program may execute the following steps:
[0271] Receiving a light field image to be recognized, where the light field image includes light sub-images at different light field angles; extracting features from each light sub-image to obtain a set of sub-aperture features corresponding to each light sub-image; fusing the multiple sets of sub-aperture features according to the object type to obtain sub-aperture fusion features under different object types; generating corresponding correlation parameters based on the correlation degree between the sub-aperture fusion features under different object types, and generating distortion parameters based on the recognition degree of the sub-aperture fusion features under different object types, where the correlation degree is generated by the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated by the difference degree between the sub-aperture fusion features under different object types; and determining the target distortion score of the light field image according to the correlation parameters and the distortion parameters.
[0272] Among them, the computer-readable storage medium may include: a read-only memory (ROM, Read Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, an optical disc, or the like.
[0273] The present application also provides a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the virtual resource allocation method provided in various alternative implementations in the foregoing embodiments.
[0274] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the image recognition methods provided in the embodiments of the present application, the beneficial effects achievable by any of the image recognition methods provided in the embodiments of the present application can be achieved. For details, see the foregoing embodiments and will not be elaborated herein.
[0275] The foregoing has introduced in detail an image recognition method, apparatus, and computer-readable storage medium provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An image recognition method, characterized in that, Including: Receiving a light field image to be recognized, where the light field image includes light sub-images at different light field angles; Performing feature extraction on each light sub-image to obtain a sub-aperture feature set corresponding to each light sub-image; Fusing the multiple sub-aperture feature sets according to object types to obtain sub-aperture fusion features under different object types; Generating corresponding correlation parameters according to the correlation degrees between the sub-aperture fusion features under different object types, and generating distortion parameters according to the recognition degrees of the sub-aperture fusion features under different object types, where the correlation degree is generated by the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated by the difference degree between the sub-aperture fusion features of different object types; Determining a target distortion score of the light field image according to the correlation parameters and the distortion parameters.
2. The method according to claim 1, wherein The fusing the multiple sub-aperture feature sets according to object types to obtain sub-aperture fusion features under different object types includes: Identifying the object type of each sub-aperture feature in each sub-aperture feature set; Obtaining the color channels of the sub-aperture features corresponding to each object type; Fusing the sub-aperture features of the same object type between the sub-aperture feature sets according to the screening rules of the color channels to obtain sub-aperture fusion features under different object types.
3. The method according to claim 2, wherein The fusing the sub-aperture features of the same object type between the sub-aperture feature sets according to the screening rules of the color channels to obtain sub-aperture fusion features under different object types includes: Fusing the sub-aperture features with the same color channels between the sub-aperture feature sets to obtain first sub-aperture fusion features corresponding to different color channels; Filtering the first sub-aperture fusion features to obtain filtered second sub-aperture fusion features; Performing feature fusion on the second sub-aperture fusion features of different color channels to obtain sub-aperture fusion features under different object types.
4. The method according to claim 1, wherein The generating corresponding correlation parameters according to the correlation degrees between the sub-aperture fusion features under different object types includes: Inputting the sub-aperture fusion features of different object types into a global context awareness module in a trained target model; Outputting correlation parameters through the global context awareness module, where the correlation parameters are generated by the global context awareness module according to the correlation degrees between the sub-aperture fusion features of different object types.
5. The method according to claim 4, wherein Before receiving the light field image to be recognized, it further includes: Obtaining a sample light field image set, where the sample light field image set includes multiple distorted sample light field images and a sample distortion score corresponding to each sample light field image; Performing joint pre-training on a preset model according to the sample light field images and the corresponding sample distortion scores to obtain a pre-trained model; Obtaining a real data set, where the real data set includes light field images with different compression ratios and a real distortion score corresponding to each light field image; Performing joint training on the pre-trained model according to the light field images with different compression ratios and the corresponding real distortion scores to obtain a target model; Then the receiving the light field image to be recognized includes: Receiving the light field image to be recognized through the target model.
6. The method according to claim 5, characterized in that, Jointly training the pre-trained model based on the light field images with different compression ratios and the corresponding true distortion scores to obtain a target model, including: Inputting the light field images with different compression ratios into the pre-trained model to obtain corresponding predicted distortion scores; Obtaining the difference value between the predicted distortion score and the true distortion score; Iteratively training the network parameters of the pre-trained model based on the difference value until the difference value converges to obtain the trained target model.
7. The method according to claim 5, characterized in that, Before generating the corresponding correlation parameters according to the correlation degree between the sub-aperture fusion features under different object types and generating the distortion parameters according to the recognition degree of the sub-aperture fusion features under different object types, further including: Obtaining the number of sub-aperture features included in the sub-aperture fusion feature corresponding to each object type; Filtering the sub-aperture fusion features with the number of sub-aperture features less than the preset feature number threshold to obtain filtered sub-aperture fusion features; Performing downsampling processing on the filtered sub-aperture fusion features to obtain target sub-aperture fusion features; Then generating the corresponding correlation parameters according to the correlation degree between the sub-aperture fusion features under different object types and generating the distortion parameters according to the recognition degree of the sub-aperture fusion features under different object types, including: Generating the corresponding correlation parameters according to the correlation degree between the target sub-aperture fusion features of different object types and generating the distortion parameters according to the recognition degree of the target sub-aperture fusion features of different object types.
8. An image recognition device, characterized in that, Including: A receiving unit, configured to receive a light field image to be recognized, where the light field image includes light sub-images with different light field angles; An extracting unit, configured to extract features from each light sub-image to obtain a set of sub-aperture features corresponding to each light sub-image; A fusing unit, configured to fuse the sets of sub-aperture features according to object types to obtain sub-aperture fusion features under different object types; A generating unit, configured to generate corresponding correlation parameters according to the correlation degree between the sub-aperture fusion features under different object types and generate distortion parameters according to the recognition degree of the sub-aperture fusion features under different object types, where the correlation degree is generated by the correlation between the sub-aperture fusion features of different object types, and the recognition degree is generated by the difference degree between the sub-aperture fusion features of different object types; A determining unit, configured to determine the target distortion score of the light field image according to the correlation parameters and the distortion parameters.
9. The device according to claim 8, wherein The fusing unit includes: An identifying sub-unit, configured to identify the object type of each sub-aperture feature in each set of sub-aperture features; An obtaining sub-unit, configured to obtain the color channels of the sub-aperture features corresponding to each object type; A fusing sub-unit, configured to fuse the sub-aperture features of the same object type between the sets of sub-aperture features according to the color channel screening rule to obtain sub-aperture fusion features under different object types.
10. A computer-readable storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-7.
11. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method according to any one of claims 1-7.
12. A computer program product, characterized in that, A computer program product includes computer instructions stored in a computer-readable storage medium. The processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to perform the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Light field image quality evaluation method based on polar plane multi-scale Gabor feature similarity
CN110310269A
Multi-image encryption method and decryption method based on light field sub-aperture image
CN110312055A