Image generation method based on depth perception circulation attention and related equipment thereof

Through the image generation method of depth-aware recurrent attention, the image depth information is integrated, which solves the problems of simulated image robustness and poor layout results in the existing technology, and achieves high-precision and high-robustness image generation, which is applied to medical imaging and vehicle claims scenarios.

CN120707670APending Publication Date: 2025-09-26PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510724309.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

Existing image generation models are unable to explicitly integrate image depth information, resulting in poor robustness of simulated images and poor image layout results.

Method used

An image generation method based on depth-aware recurrent attention is adopted. By receiving the original depth data, the image preprocessing layer is used for preprocessing, the deep feature encoding layer is used for feature encoding, the deep feature layout processing layer is used for hierarchical attention calculation, and the visual rendering layer is used to generate a visual image.

Benefits of technology

The image generation model fully integrates information at different depth levels, improving the robustness and accuracy of simulated images, making it suitable for accurate analysis of medical imaging and vehicle accident claims scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120707670A_ABST
    Figure CN120707670A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, can be applied to financial science and technology or medical health related business scenes, and relates to an image generation method based on depth perception circulation attention and related equipment thereof. The image generation model comprises an image preprocessing layer, a depth feature coding layer, a depth feature layout processing layer and a visual rendering layer. The depth feature layout processing layer is adopted to carry out hierarchical attention calculation on depth level feature embedded representation, and a data layout result is generated according to a hierarchical attention calculation result, so that image information of different depth levels is more fully integrated; and high robustness and high accuracy of the image finally generated by the image generation model are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and is applied to image simulation generation scenarios using neural networks, and relates to an image generation method based on depth perception cyclic attention and related equipment. Background Art

[0002] Currently, with the rapid development of deep learning and computer vision technologies, it is now possible to extract the required key information directly from images through high-precision semantic segmentation of image floor plans. This method can identify different regions and structural elements in the image floor plan, such as individual objects, thereby providing spatial layout information for image simulation modeling.

[0003] In the existing technology, most image generation models are still unable to explicitly integrate image depth information, resulting in poor ability to model the spatial hierarchical relationship of images. The introduction of existing multi-scale feature extraction methods, such as the FPN (Feature Pyramid Network) method, although able to extract features of different resolutions from images, still lacks the corresponding feature fusion interaction mechanism, resulting in the loss of key details or the introduction of noise during the feature fusion process, which ultimately easily leads to problems with the robustness of the generated simulated images and poor image layout results. Summary of the Invention

[0004] The purpose of the embodiments of the present application is to propose an image generation method based on depth-perceived recurrent attention and related equipment to solve the problem that the robustness of the generated simulated image and the image layout results are poor during the existing image generation.

[0005] In a first aspect, the embodiments of the present application provide an image generation method based on depth-aware recurrent attention, which adopts the following technical solutions:

[0006] The image generation method based on depth-aware recurrent attention includes the following steps:

[0007] Receive raw depth data, wherein the raw depth data includes raw image data, the raw image data includes a depth value range of the image, a target object label in the image, and a depth value range corresponding to each target object in the image, wherein the depth value represents a spatial distance value between an image capture point and a target object in the image;

[0008] Inputting the original depth data into a pre-trained image generation model, wherein the image generation model includes an image pre-processing layer, a depth feature encoding layer, a depth feature layout processing layer, and a visualization rendering layer;

[0009] Preprocessing the original depth data using a normalization layer and a size standard preset in the image preprocessing layer to obtain a preliminary depth map;

[0010] Performing feature encoding on the preliminary depth map using the depth feature encoding layer to generate a depth-level feature embedding representation;

[0011] Using the deep feature layout processing layer to perform hierarchical attention calculation on the deep level feature embedding representation, and generating a data layout result according to the hierarchical attention calculation result;

[0012] The data layout result is subjected to visual rendering conversion by the visual rendering layer to generate a visual image.

[0013] In a second aspect, the embodiments of the present application also provide an image generation device based on depth-aware cyclic attention, which adopts the following technical solution:

[0014] An image generation device based on depth-aware recurrent attention, comprising:

[0015] A data receiving module, configured to receive raw depth data, wherein the raw depth data includes raw image data, the raw image data including a depth value range of the image, a target object label in the image, and a depth value range corresponding to each target object in the image, wherein the depth value represents a spatial distance value between an image capture point and a target object in the image;

[0016] A data input module is used to input the original depth data into a pre-trained image generation model, wherein the image generation model includes an image preprocessing layer, a depth feature encoding layer, a depth feature layout processing layer, and a visualization rendering layer;

[0017] A data preprocessing module, configured to preprocess the original depth data using a normalization layer and size standard preset in the image preprocessing layer to obtain a preliminary depth map;

[0018] a depth feature encoding module, configured to perform feature encoding on the preliminary depth map using the depth feature encoding layer to generate a depth-level feature embedding representation;

[0019] a data layout generation module, configured to perform a hierarchical attention calculation on the depth-level feature embedding representation using the depth feature layout processing layer, and generate a data layout result according to the hierarchical attention calculation result;

[0020] The visual rendering module is used to perform visual rendering conversion on the data layout result through the visual rendering layer to generate a visual image.

[0021] In a third aspect, an embodiment of the present application further provides a computer device that adopts the following technical solution:

[0022] A computer device includes a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, it implements the steps of the above-mentioned image generation method based on depth perception recurrent attention.

[0023] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0024] A computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the image generation method based on depth-aware recurrent attention as described above.

[0025] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0026] The image generation method based on depth-aware recurrent attention described in the embodiment of the present application receives raw depth data, wherein the raw depth data includes raw image data; inputs the raw depth data into a pre-trained image generation model, wherein the image generation model includes an image preprocessing layer, a depth feature encoding layer, a depth feature layout processing layer, and a visualization rendering layer; preprocesses the raw depth data through a normalization layer and a size standard preset in the image preprocessing layer to obtain a preliminary depth map; uses the depth feature encoding layer to feature encode the preliminary depth map to generate a depth-level feature embedding representation; uses the depth feature layout processing layer to perform hierarchical attention calculation on the depth-level feature embedding representation, and generates a data layout result based on the hierarchical attention calculation result; and converts the data layout result into a visualization rendering conversion through the visualization rendering layer to generate a visualization image. By using the depth feature layout processing layer to perform hierarchical attention calculation on the depth-level feature embedding representation, and generating a data layout result based on the hierarchical attention calculation result, the image information of different depth levels is more fully integrated, ensuring the high robustness and high accuracy of the image finally generated by the image generation model. Applying the image generation method to medical imaging technology can assist medical personnel in more accurately performing image analysis in combination with image simulation results, or applying the image generation method to vehicle traffic accident claims scenarios can assist the claims business end in more accurately identifying vehicle damage conditions. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the solutions in this application, a brief introduction will be given below to the drawings required for use in the description of the embodiments of this application. Obviously, the drawings described below are some embodiments of this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0028] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0029] Figure 2 is a flowchart of an embodiment of a method for generating an image based on depth-aware recurrent attention according to the present application;

[0030] Figure 3 yes Figure 2 A flowchart of a specific embodiment of step 204 is shown;

[0031] Figure 4 yes Figure 2 A flowchart of a specific embodiment of step 205 is shown;

[0032] Figure 5 yes Figure 4 A flowchart of a specific embodiment of step 401 is shown;

[0033] Figure 6 yes Figure 2 A flowchart of a specific embodiment of step 206 is shown;

[0034] Figure 7 is a structural diagram of an embodiment of an image generation device based on depth-aware cyclic attention according to the present application;

[0035] Figure 8 It is a structural diagram of an embodiment of a computer device according to the present application. DETAILED DESCRIPTION

[0036] Unless otherwise defined, all technical and scientific terms used herein have the same meanings as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of the application are for the purpose of describing specific embodiments only and are not intended to limit this application. The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned drawings are used to distinguish different objects, not to describe a specific order.

[0037] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0038] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0039] like Figure 1 As shown, system architecture 100 may include a terminal device 101, a network 102, and a server 103. Terminal device 101 may be a laptop computer 1011, a tablet computer 1012, or a mobile phone 1013. Network 102 is a medium for providing a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables.

[0040] The user can use the terminal device 101 to interact with the server 103 via the network 102 to receive or send messages, etc. Various communication client applications can be installed on the terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.

[0041] The terminal device 101 can be various electronic devices with a display screen and supporting web browsing. In addition to the laptop computer 1011, tablet computer 1012 or mobile phone 1013, the terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer and a desktop computer, etc.

[0042] The server 103 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal device 101 .

[0043] It should be noted that the image generation method based on depth perception cyclic attention provided in the embodiment of the present application is generally executed by a server, and accordingly, the image generation device based on depth perception cyclic attention is generally set in the server.

[0044] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0045] Continue to refer Figure 2 , shows a flow chart of an embodiment of a method for generating an image based on depth-aware recurrent attention according to the present application. The method for generating an image based on depth-aware recurrent attention comprises the following steps:

[0046] Step 201: Receive original depth data.

[0047] The original depth data includes original image data, the original image data includes the depth value range of the image, the target object label in the image, and the depth value range corresponding to the target object in the image. The depth value includes the spatial distance value between the image shooting point and the target object in the image;

[0048] It should be understood that the raw depth data refers to data obtained directly through measurement or observation without any processing or manipulation. For example, the raw depth image refers to an image obtained directly through shooting and without any processing or manipulation. Accordingly, the raw depth image, for example, in a financial business scenario, is business image data directly scanned, such as a scan of a personal ID card (ID card, bank card, social security card, etc.) uploaded by a user when handling financial business; another example is in a car insurance claim business scenario, an image of a damaged vehicle directly taken and received by the claims processing end; another example is in a medical image scanning scenario, an image of a lesion or an X-ray diagnostic image obtained by directly scanning a medical patient.

[0049] In this embodiment, the depth value also includes the size value of the target object in the image;

[0050] In this embodiment, the depth value also includes the object appearance color range value of the target object and background objects in the image, for example: the RGB color value corresponding to the current target object, specifically, the R, G, and B values ​​are all in the [0,255] color value range.

[0051] It should be understood that the depth value refers to the multi-dimensional feature data of the target object that can be extracted from the image, and is not limited to the three example feature data: the spatial distance value between the image shooting point and the target object in the image, the size value of the target object in the image, and the object appearance color range value of the target object and the background object in the image.

[0052] By receiving the original depth data, when the neural network is used for subsequent image simulation generation, the spatial distance values, color values ​​and object size values ​​between objects in the original image can be combined to generate a more realistic simulated image.

[0053] Step 202: input the original depth data into the pre-trained image generation model.

[0054] The image generation model includes an image preprocessing layer, a depth feature encoding layer, a depth feature layout processing layer, and a visualization rendering layer;

[0055] By inputting raw depth data into a pre-trained image generation model, simulated image generation using the image generation model is more automated. Furthermore, by setting or marking key areas of interest or parts of interest, simulation generation can be targeted to only those key target areas or objects of interest in the image, eliminating non-key areas from the generated simulated image, allowing for easier use of the information conveyed by the image for business processing.

[0056] Specifically, for example: in a vehicle claims scenario, focus on the collision damage area in the uploaded vehicle image and simulate and generate this area; or in a medical image detection scenario, focus on the area corresponding to the diseased organ and simulate and generate this area, thereby more intuitively assisting relevant business processing personnel in business analysis.

[0057] Step 203 : Preprocess the original depth data using the normalization layer and size standard preset in the image preprocessing layer to obtain a preliminary depth map.

[0058] Wherein, the preprocessing includes performing data normalization, size standardization and initial convolution processing on the image data;

[0059] It should be understood that the normalization refers to adjusting the pixel values ​​of the image to a specific range, usually mapping the pixel values ​​to the interval [0,1] or [-1,1]. The purpose of normalization is to eliminate the scale differences between different features, making the model easier to train. For example: the minimum-maximum normalization processing method can linearly map the data to the interval [0,1], or the Z-score normalization processing method converts the data into a normal distribution with a mean of 0 and a standard deviation of 1. The size standardization refers to adjusting the size of the image to a fixed resolution, usually common sizes such as 224x224, 299x299, etc. The purpose of size standardization is to ensure that the input image has a consistent size for easy model processing.

[0060] In this embodiment, the original depth data is preprocessed using the normalization layer and size standard preset in the image preprocessing layer to obtain preprocessed data. Generally speaking, this is to facilitate subsequent processing of the image generation model.

[0061] The initial convolution process includes performing an initial image feature extraction using a convolution layer for image feature extraction. After obtaining the initial image feature extraction results, a corresponding intermediate image can be directly generated based on the initial image feature extraction results, and the intermediate image can be used as the initial depth map. Alternatively, the intermediate image can be superimposed with the original depth image to obtain a superimposed image as the initial depth map, so that the initial depth map can be directly used to generate image simulation.

[0062] Step 204: perform feature encoding on the preliminary depth map using the depth feature encoding layer to generate a depth-level feature embedding representation.

[0063] In this embodiment, the deep feature coding layer is composed of a pre-trained encoder based on the ResNet-50 processing network. The pre-trained encoder based on the ResNet-50 processing network is used to perform feature encoding on the preliminary depth map to generate a deep-level feature embedding representation, replacing the traditional multi-head self-attention layer for feature encoding. There is no need for manual feature engineering, which saves a lot of manpower consumption and avoids the errors generated during manual feature engineering, so that the generated deep-level feature embedding representation is not subject to the subjective influence of human factors.

[0064] Step 205: Use the deep feature layout processing layer to perform hierarchical attention calculation on the deep level feature embedding representation, and generate a data layout result according to the hierarchical attention calculation result.

[0065] Among them, the deep feature layout processing layer is composed of a cyclic attention calculation network based on depth perception, a fully connected attention integration network and a data layout representation generation network. In essence, the cyclic attention calculation network adopts a cyclic calculation method to calculate the attention weight of the image depth feature, so that it can combine the cyclic processing method to obtain the most comprehensive and important image depth features. Afterwards, the fully connected attention integration network is used to perform fully connected integration processing on all the acquired image depth features. Finally, the integrated attention output features are layout output through the data layout representation generation network to obtain a preliminary data layout representation as the data layout result.

[0066] Specifically, the depth-perception-based recurrent attention computing network introduces a depth-perception gating mechanism, an attention head cyclic scheduling mechanism and a hierarchical attention combination mechanism. The depth-perception gating mechanism dynamically allocates attention weights of different depth levels to all depth-level feature embedding representations according to the difference in input depth; the attention head cyclic scheduling mechanism uses a cyclic activation method to periodically activate the attention of different depth levels to obtain the attention activation results corresponding to all attention heads; the hierarchical attention combination mechanism is used to integrate the attention activation results corresponding to all attention heads.

[0067] In this embodiment, by adopting the depth feature layout processing layer to perform hierarchical attention calculation on the depth level feature embedding representation, and generating data layout results based on the hierarchical attention calculation results, it is achieved that under the combination of the depth perception gating mechanism, the attention head loop scheduling mechanism and the hierarchical attention combination mechanism, the attention weight calculation of image features at different depth levels is cyclically performed, which fully ensures the comprehensiveness of the acquisition of image depth features.

[0068] Step 206: Perform visual rendering conversion on the data layout result through the visual rendering layer to generate a visual image.

[0069] In this embodiment, by receiving original depth data, wherein the original depth data includes original image data; inputting the original depth data into a pre-trained image generation model, wherein the image generation model includes an image preprocessing layer, a depth feature encoding layer, a depth feature layout processing layer and a visual rendering layer; preprocessing the original depth data through the normalization layer and size standard preset in the image preprocessing layer to obtain a preliminary depth map; using the depth feature encoding layer to feature encode the preliminary depth map to generate a depth-level feature embedding representation; using the depth feature layout processing layer to perform hierarchical attention calculation on the depth-level feature embedding representation, and generating a data layout result based on the hierarchical attention calculation result; the data layout result is converted into a visual rendering through the visual rendering layer to generate a visual image. Applying the image generation method to medical imaging technology can assist medical personnel in more accurately performing image analysis in combination with image simulation results, or applying the image generation method to vehicle traffic accident claims scenarios can assist the claims business end in more accurately identifying vehicle damage conditions.

[0070] In this embodiment, the step of receiving raw depth data specifically includes receiving raw depth data transmitted by a target depth sensor via a preset depth sensor interface. The target depth sensor is, for example, a high-magnification camera lens. The preset depth sensor interface is, for example, an image upload interface of the high-magnification camera lens that is responsible for sending the captured image to the target processor.

[0071] Continue to refer Figure 3 , Figure 3 yes Figure 2 The flowchart of a specific embodiment of step 204 shown includes the following steps:

[0072] Step 301: input the preliminary depth map into the pre-trained encoder based on the ResNet-50 processing network;

[0073] Step 302: Use the pre-trained encoder based on the ResNet-50 processing network to perform feature encoding on the preliminary depth map to generate a depth-level feature embedding representation corresponding to the original depth data.

[0074] In this embodiment, the preliminary depth map is input into the pre-trained encoder based on the ResNet-50 processing network, and the pre-trained encoder based on the ResNet-50 processing network is used to perform feature encoding on the preliminary depth map to generate a depth-level feature embedding representation corresponding to the original depth data. Compared with the previous feature encoding using the traditional multi-head self-attention layer, there is no need for manual feature engineering, which saves a lot of manpower consumption and avoids the errors generated during manual feature engineering, so that the generated depth-level feature embedding representation is not subject to the subjective influence of human factors.

[0075] Continue to refer Figure 4 , Figure 4 yes Figure 2 The flowchart of a specific embodiment of step 205 shown includes the following steps:

[0076] Step 401, using the depth-aware recurrent attention calculation network to perform hierarchical attention calculation on the depth-level feature embedding representation to generate hierarchical attention weights;

[0077] Step 402: integrating the hierarchical attention weights through the fully connected attention integration network to obtain integrated attention output features;

[0078] Step 403: The integrated attention output features are output through the data layout representation generation network to obtain a preliminary data layout representation as the data layout result.

[0079] The data layout result is used to characterize the relative position relationship between target objects in the image, the spatial size and spatial position information of the target objects in the image.

[0080] In this embodiment, the cyclic attention calculation network adopts a cyclic calculation method to calculate the attention weight of the image depth feature, so that it can combine the cyclic processing method to obtain the most comprehensive and important image depth features; then, the fully connected attention integration network is used to perform fully connected integration processing on all the acquired image depth features, and finally, the integrated attention output features are layout output through the data layout representation generation network to obtain a preliminary data layout representation as the data layout result, so as to provide support for subsequent image layout rendering.

[0081] Continue to refer Figure 5 , Figure 5 yes Figure 4 The flowchart of a specific embodiment of step 401 includes the following steps:

[0082] Step 501, using the depth-aware gating mechanism in the recurrent attention computation network to assign attention weights to all depth-level feature embedding representations;

[0083] The depth-aware gating mechanism dynamically assigns attention weights of different depth levels to all depth-level feature embedding representations based on the difference in input depth.

[0084] Step 502: Based on the attention head cyclic scheduling mechanism and the depth levels of the attention weights corresponding to all depth-level feature embedding representations, a cyclic activation method is used to periodically activate the attention of different depth levels to obtain attention activation results corresponding to all attention heads.

[0085] Specifically, the step of periodically activating the attention at different depth levels using a cyclic activation method according to the attention head cyclic scheduling mechanism and the depth level of the attention weight corresponding to all depth-level feature embedding representations to obtain the attention activation results corresponding to all attention heads includes: according to a preset activation algorithm formula:

[0086]

[0087] Periodically adjust the activation frequency of all attention heads, where p i (t) represents the activation probability of the i-th attention head when the number of loop activations is t, σ(*) represents the sigmoid activation function, and “*” in the current formula represents T represents the preset cycle period, φ iis the phase offset, α is the scaling parameter, and β is the bias parameter; then, the attention at different depth levels is activated according to the activation frequencies corresponding to different attention heads, and the attention activation results corresponding to all attention heads are obtained.

[0088] By using the above-mentioned cyclic activation method to periodically activate the attention of different depth levels, it is possible to explicitly incorporate image depth information into the design of the attention mechanism. Through dynamic routing and periodic scheduling, the model can adaptively focus on features at different depth levels, thereby ensuring the subsequent generation of a more reasonable and accurate layout structure. Each attention head corresponds to a specific depth level.

[0089] Through the synergy of the depth perception gating mechanism and the attention head loop scheduling mechanism, the hierarchical attention combination output is finally achieved under the combination of the hierarchical attention combination mechanism.

[0090] In step 503, the hierarchical attention combination mechanism is used to integrate the attention activation results corresponding to all attention heads to obtain the integrated hierarchical attention weights.

[0091] Specifically, the hierarchical attention combination mechanism includes: after obtaining the gating weight calculated by the depth perception gating mechanism, performing a dot multiplication operation with the activation frequency of all attention heads output by the attention head loop scheduling mechanism to obtain the final output hierarchical attention combination, that is, assuming that the gating weight calculated by the depth perception gating mechanism is G i , the activation frequency of all attention heads output by the attention head loop scheduling mechanism is P i (t), where i is the number of the attention head, each attention head corresponds to a different depth level, and each depth level corresponds to a different depth layer. The hierarchical attention combination mechanism adopts ω i =G i ·P i (t) method to calculate the hierarchical attention combinations corresponding to different depth levels, where “·” represents the operation method, such as: dot product operation.

[0092] In this embodiment, by combining the depth perception gating mechanism, the attention head loop scheduling mechanism and the hierarchical attention combination mechanism, the attention weights of image features at different depth levels are calculated cyclically, fully ensuring the comprehensive acquisition and activation of image depth features, and integrating hierarchical attention weights.

[0093] Continue to refer Figure 6 , Figure 6 yes Figure 2 The flowchart of a specific embodiment of step 206 includes the following steps:

[0094] Step 601: Obtain, through the visualization rendering layer, from the data layout result, attention output features corresponding to the corresponding attention heads in sequence according to a depth hierarchy from deep to shallow, wherein the depth hierarchy from deep to shallow corresponds to a distance from far to near in the target image space;

[0095] Step 602: Visually render the target objects in the target image in order of spatial distance from far to near based on the attention output features corresponding to all attention heads;

[0096] Step 603 : until all target objects in the target image are visually rendered, the visual image is obtained.

[0097] Specifically, when performing visual rendering, the attention output features corresponding to the corresponding attention heads are first obtained from the data layout result in sequence according to the depth level from deep to shallow, and the visual rendering of the target objects with spatial distances from far to near in the target image is completed, that is, the image rendering of the object with the farthest spatial distance is first performed, and then the progressive rendering is gradually performed to finally generate the visual image. It should be understood that when the image simulation is generated, the farther the spatial distance of the object is, the lower the resolution of the corresponding image features is, that is, the farther the distance is, the blurrier it is, and the closer the distance is, the clearer the simulation is.

[0098] In this embodiment, by receiving original depth data, wherein the original depth data includes original image data; inputting the original depth data into a pre-trained image generation model, wherein the image generation model includes an image preprocessing layer, a depth feature encoding layer, a depth feature layout processing layer and a visual rendering layer; preprocessing the original depth data through the normalization layer and size standard preset in the image preprocessing layer to obtain a preliminary depth map; using the depth feature encoding layer to feature encode the preliminary depth map to generate a depth-level feature embedding representation; using the depth feature layout processing layer to perform hierarchical attention calculation on the depth-level feature embedding representation, and generating a data layout result based on the hierarchical attention calculation result; the data layout result is converted into a visual rendering through the visual rendering layer to generate a visual image. Applying the image generation method to medical imaging technology can assist medical personnel in more accurately performing image analysis in combination with image simulation results, or applying the image generation method to vehicle traffic accident claims scenarios can assist the claims business end in more accurately identifying vehicle damage conditions.

[0099] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.

[0100] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0101] In this embodiment, by receiving original depth data, wherein the original depth data includes original image data; inputting the original depth data into a pre-trained image generation model, wherein the image generation model includes an image preprocessing layer, a depth feature encoding layer, a depth feature layout processing layer and a visual rendering layer; preprocessing the original depth data through the normalization layer and size standard preset in the image preprocessing layer to obtain a preliminary depth map; using the depth feature encoding layer to feature encode the preliminary depth map to generate a depth-level feature embedding representation; using the depth feature layout processing layer to perform hierarchical attention calculation on the depth-level feature embedding representation, and generating a data layout result based on the hierarchical attention calculation result; the data layout result is converted into a visual rendering through the visual rendering layer to generate a visual image. Applying the image generation method to medical imaging technology can assist medical personnel in more accurately performing image analysis in combination with image simulation results, or applying the image generation method to vehicle traffic accident claims scenarios can assist the claims business end in more accurately identifying vehicle damage conditions.

[0102] Further references Figure 7 , as a response to the above Figure 2 The present application provides an embodiment of an image generation device based on depth perception cyclic attention. Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0103] like Figure 7 As shown, the image generation device 700 based on depth perception cyclic attention described in this embodiment includes: a data receiving module 701, a data input module 702, a data preprocessing module 703, a depth feature encoding module 704, a data layout generation module 705 and a visualization rendering module 706.

[0104] in:

[0105] A data receiving module 701 is configured to receive raw depth data, wherein the raw depth data includes raw image data, the raw image data including a depth value range of the image, a target object label in the image, and a depth value range corresponding to each target object in the image, wherein the depth value represents a spatial distance value between an image capture point and a target object in the image;

[0106] A data input module 702 is configured to input the raw depth data into a pre-trained image generation model, wherein the image generation model includes an image pre-processing layer, a depth feature encoding layer, a depth feature layout processing layer, and a visualization rendering layer;

[0107] A data preprocessing module 703 is configured to preprocess the original depth data using a normalization layer and size standard preset in the image preprocessing layer to obtain a preliminary depth map;

[0108] a depth feature encoding module 704 for encoding the preliminary depth map using the depth feature encoding layer to generate a depth-level feature embedding representation;

[0109] A data layout generation module 705 is configured to perform a hierarchical attention calculation on the depth-level feature embedding representation using the depth feature layout processing layer, and generate a data layout result based on the hierarchical attention calculation result;

[0110] The visual rendering module 706 is configured to perform visual rendering conversion on the data layout result through the visual rendering layer to generate a visual image.

[0111] The present application receives original depth data, wherein the original depth data includes original image data; inputs the original depth data into a pre-trained image generation model, wherein the image generation model includes an image preprocessing layer, a depth feature encoding layer, a depth feature layout processing layer and a visual rendering layer; preprocesses the original depth data through the normalization layer and size standard preset in the image preprocessing layer to obtain a preliminary depth map; uses the depth feature encoding layer to feature encode the preliminary depth map to generate a depth-level feature embedding representation; uses the depth feature layout processing layer to perform hierarchical attention calculation on the depth-level feature embedding representation, and generates a data layout result based on the hierarchical attention calculation result; and converts the data layout result into a visual rendering conversion through the visual rendering layer to generate a visual image. Applying the image generation method to medical imaging technology can assist medical personnel in more accurately performing image analysis in combination with image simulation results, or applying the image generation method to vehicle traffic accident claims scenarios can assist the claims business end in more accurately identifying vehicle damage conditions.

[0112] In this embodiment, the depth feature encoding module 704 includes a preliminary depth map input unit and a depth feature encoding unit.

[0113] A preliminary depth map input unit, configured to input the preliminary depth map into the pre-trained encoder based on the ResNet-50 processing network;

[0114] A depth feature encoding unit is used to perform feature encoding on the preliminary depth map using the pre-trained encoder based on the ResNet-50 processing network to generate a depth-level feature embedding representation corresponding to the original depth data.

[0115] In this embodiment, the data layout generation module 705 includes a hierarchical attention calculation unit, a fully connected integration unit, and a data layout generation unit.

[0116] A hierarchical attention calculation unit, configured to perform hierarchical attention calculation on the depth-level feature embedding representation using the depth-aware recurrent attention calculation network to generate hierarchical attention weights;

[0117] a fully connected integration unit, configured to integrate the hierarchical attention weights through the fully connected attention integration network to obtain an integrated attention output feature;

[0118] A data layout generation unit is used to layout and output the integrated attention output features through the data layout representation generation network to obtain a preliminary data layout representation as the data layout result.

[0119] In this embodiment, the data layout generation module 705 further includes an attention weight allocation unit, an attention cycle activation unit, and an attention combination unit.

[0120] an attention weight assignment unit, configured to assign attention weights to all depth-level feature embedding representations using a depth-aware gating mechanism in the recurrent attention computation network;

[0121] An attention recurrent activation unit is configured to periodically activate attention at different depth levels using a recurrent activation method according to the attention head recurrent scheduling mechanism and the depth levels of the attention weights corresponding to all depth-level feature embedding representations, thereby obtaining attention activation results corresponding to all attention heads.

[0122] The attention combination unit is used to adopt the hierarchical attention combination mechanism to integrate the attention activation results corresponding to all attention heads to obtain the integrated hierarchical attention weights.

[0123] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware via computer-readable instructions. The computer-readable instructions can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0124] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.

[0125] To solve the above technical problems, the present application also provides a computer device. Figure 8 , Figure 8 This is a basic structural block diagram of the computer device in this embodiment.

[0126] The computer device 8 includes a memory 8a, a processor 8b, and a network interface 8c that are interconnected via a system bus. Figure 8 Only a computer device 8 having components such as a memory 8a, a processor 8b, and a network interface 8c is shown. However, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead. It should be understood by those skilled in the art that a computer device herein is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to a microprocessor, an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a digital signal processor (DSP), an embedded device, etc.

[0127] The computer device may be a desktop computer, notebook computer, PDA, cloud server, etc. The computer device may interact with the user via a keyboard, mouse, remote control, touchpad, or voice control device.

[0128] The memory 8a includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 8a can be an internal storage unit of the computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the memory 8a can also be an external storage device of the computer device 8, such as a plug-in hard disk equipped on the computer device 8, a smart memory card (SMC), a secure digital (SD) card, a flash memory card, etc. Of course, the memory 8a can also include both the internal storage unit of the computer device 8 and its external storage device. In this embodiment, the memory 8a is generally used to store the operating system and various application software installed on the computer device 8, such as computer-readable instructions for the image generation method based on depth-aware recurrent attention. In addition, the memory 8a can also be used to temporarily store various types of data that have been output or are to be output.

[0129] In some embodiments, the processor 8b may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 8b is generally used to control the overall operation of the computer device 8. In this embodiment, the processor 8b is used to execute computer-readable instructions or process data stored in the memory 8a, such as computer-readable instructions for executing the depth-aware recurrent attention-based image generation method.

[0130] The network interface 8c may include a wireless network interface or a wired network interface. The network interface 8c is generally used to establish a communication connection between the computer device 8 and other electronic devices.

[0131] The computer device proposed in this embodiment belongs to the field of artificial intelligence technology and is applied to image simulation generation scenarios using neural networks. The present application receives raw depth data, wherein the raw depth data includes raw image data; inputs the raw depth data into a pre-trained image generation model, wherein the image generation model includes an image preprocessing layer, a depth feature encoding layer, a depth feature layout processing layer, and a visualization rendering layer; preprocesses the raw depth data through a normalization layer and a size standard preset in the image preprocessing layer to obtain a preliminary depth map; uses the depth feature encoding layer to feature encode the preliminary depth map to generate a depth-level feature embedding representation; uses the depth feature layout processing layer to perform hierarchical attention calculation on the depth-level feature embedding representation, and generates a data layout result based on the hierarchical attention calculation result; and converts the data layout result into a visualization rendering conversion through the visualization rendering layer to generate a visualization image. Applying the image generation method to medical imaging technology can assist medical personnel in more accurately performing image analysis in combination with image simulation results, or applying the image generation method to a vehicle traffic accident claims scenario can assist the claims business end in more accurately identifying vehicle damage conditions.

[0132] The present application also provides another embodiment, namely, providing a computer-readable storage medium, which stores computer-readable instructions, and the computer-readable instructions can be executed by a processor to enable the processor to perform the steps of the image generation method based on depth perception cyclic attention as described above.

[0133] The computer-readable storage medium proposed in this embodiment belongs to the field of artificial intelligence technology and is applied to image simulation generation scenarios using neural networks. This application receives raw depth data, wherein the raw depth data includes raw image data; inputs the raw depth data into a pre-trained image generation model, wherein the image generation model includes an image preprocessing layer, a depth feature encoding layer, a depth feature layout processing layer, and a visualization rendering layer; preprocesses the raw depth data through a normalization layer and a size standard preset in the image preprocessing layer to obtain a preliminary depth map; uses the depth feature encoding layer to feature encode the preliminary depth map to generate a depth-level feature embedding representation; uses the depth feature layout processing layer to perform hierarchical attention calculation on the depth-level feature embedding representation, and generates a data layout result based on the hierarchical attention calculation result; and converts the data layout result into a visualization rendering conversion through the visualization rendering layer to generate a visualization image. Applying the image generation method to medical imaging technology can assist medical personnel in performing image analysis more accurately in combination with image simulation results, or applying the image generation method to vehicle traffic accident claims scenarios can assist the claims business end in more accurately identifying vehicle damage conditions.

[0134] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0135] Obviously, the embodiments described above are only some of the embodiments of the present application, rather than all of the embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application has been described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific embodiments, or to make equivalent replacements for some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is also within the scope of patent protection of this application. The non-company software tools or components that appear in the embodiments of this application are merely examples and do not represent actual use.

Claims

1. Image generation method based on depth-aware recurrent attention, characterized by: The steps include: Receive raw depth data, wherein the raw depth data includes raw image data, the raw image data includes a depth value range of the image, a target object label in the image, and a depth value range corresponding to each target object in the image, wherein the depth value represents a spatial distance value between an image capture point and a target object in the image; Inputting the original depth data into a pre-trained image generation model, wherein the image generation model includes an image pre-processing layer, a depth feature encoding layer, a depth feature layout processing layer, and a visualization rendering layer; Preprocessing the original depth data using a normalization layer and a size standard preset in the image preprocessing layer to obtain a preliminary depth map; Performing feature encoding on the preliminary depth map using the depth feature encoding layer to generate a depth-level feature embedding representation; Using the deep feature layout processing layer to perform hierarchical attention calculation on the deep level feature embedding representation, and generating a data layout result according to the hierarchical attention calculation result; The data layout result is subjected to visual rendering conversion by the visual rendering layer to generate a visual image.

2. The image generation method based on depth-aware recurrent attention according to claim 1, characterized in that The step of receiving the original depth data includes: receiving the original depth data transmitted by the target depth sensor through a preset depth sensor interface.

3. The image generation method based on depth-aware recurrent attention according to claim 1, characterized in that The depth feature coding layer is composed of a pre-trained encoder based on a ResNet-50 processing network. The step of using the depth feature coding layer to perform feature encoding on the preliminary depth map to generate a depth-level feature embedding representation specifically includes: Inputting the preliminary depth map into the pre-trained encoder based on the ResNet-50 processing network; The preliminary depth map is feature encoded using the pre-trained encoder based on the ResNet-50 processing network to generate a depth-level feature embedding representation corresponding to the original depth data.

4. The image generation method based on depth-aware recurrent attention according to claim 1, characterized in that The deep feature layout processing layer is composed of a recurrent attention calculation network based on depth perception, a fully connected attention integration network, and a data layout representation generation network. The step of using the deep feature layout processing layer to perform hierarchical attention calculation on the deep level feature embedding representation and generating a data layout result based on the hierarchical attention calculation result specifically includes: Using the depth-aware recurrent attention calculation network to perform hierarchical attention calculation on the depth-level feature embedding representation to generate hierarchical attention weights; The hierarchical attention weights are integrated through the fully connected attention integration network to obtain integrated attention output features; The integrated attention output features are layout-outputted by the data layout representation generation network to obtain a preliminary data layout representation as the data layout result, wherein the data layout result is used to characterize the relative position relationship between target objects in the image, the spatial size and spatial position information of the target objects in the image.

5. The image generation method based on depth-aware recurrent attention according to claim 4, characterized in that The depth-aware recurrent attention calculation network includes a depth-aware gating mechanism, an attention head recurrent scheduling mechanism, and a hierarchical attention combination mechanism. The step of using the depth-aware recurrent attention calculation network to perform hierarchical attention calculation on the depth-level feature embedding representation to generate hierarchical attention weights specifically includes: Utilizing a depth-aware gating mechanism in the recurrent attention computation network to assign attention weights to all depth-level feature embedding representations, wherein the depth-aware gating mechanism dynamically assigns attention weights of different depth levels to all depth-level feature embedding representations based on differences in input depth; According to the attention head cyclic scheduling mechanism and the depth level of the attention weight corresponding to all depth-level feature embedding representations, a cyclic activation method is used to periodically activate the attention of different depth levels to obtain the attention activation results corresponding to all attention heads; The hierarchical attention combination mechanism is adopted to integrate the attention activation results corresponding to all attention heads to obtain the integrated hierarchical attention weights.

6. The image generation method based on depth-aware recurrent attention according to claim 5, characterized in that The step of periodically activating the attention at different depth levels using a cyclic activation method according to the attention head cyclic scheduling mechanism and the depth levels of the attention weights corresponding to all depth-level feature embedding representations to obtain the attention activation results corresponding to all attention heads includes: According to the preset activation algorithm formula: Periodically adjust the activation frequency of all attention heads, where p i (t) represents the activation probability of the i-th attention head when the number of loop activations is t, σ(*) represents the sigmoid activation function, and "*" in the current formula represents T represents the preset cycle period, φ i is the phase offset, α is the scaling parameter, and β is the bias parameter; Attention at different depth levels is activated according to the activation frequencies corresponding to different attention heads, and the attention activation results corresponding to all attention heads are obtained.

7. The image generation method based on depth-aware recurrent attention according to any one of claims 1 to 6, characterized in that: The step of performing visual rendering conversion on the data layout result through the visual rendering layer to generate a visual image specifically includes: Obtaining, through the visualization rendering layer, from the data layout result in a depth-to-shallow manner, the attention output features corresponding to the corresponding attention heads, wherein the depth-to-shallow manner corresponds to a far-to-near manner in the spatial distance of the target image; According to the attention output features corresponding to all attention heads, target objects in the target image are sequentially visualized and rendered from far to near spatial distances; The visual rendering of all target objects in the target image is completed to obtain the visual image.

8. An image generation device based on depth-aware recurrent attention, characterized in that: The image generation device based on depth perception cyclic attention is used to implement the steps of the image generation method based on depth perception cyclic attention according to any one of claims 1 to 7, and the image generation device based on depth perception cyclic attention comprises: A data receiving module, configured to receive raw depth data, wherein the raw depth data includes raw image data, the raw image data including a depth value range of the image, a target object label in the image, and a depth value range corresponding to each target object in the image, wherein the depth value represents a spatial distance value between an image capture point and a target object in the image; A data input module is used to input the original depth data into a pre-trained image generation model, wherein the image generation model includes an image preprocessing layer, a depth feature encoding layer, a depth feature layout processing layer, and a visualization rendering layer; A data preprocessing module, configured to preprocess the original depth data using a normalization layer and size standard preset in the image preprocessing layer to obtain a preliminary depth map; a depth feature encoding module, configured to perform feature encoding on the preliminary depth map using the depth feature encoding layer to generate a depth-level feature embedding representation; a data layout generation module, configured to perform a hierarchical attention calculation on the depth-level feature embedding representation using the depth feature layout processing layer, and generate a data layout result according to the hierarchical attention calculation result; The visual rendering module is used to perform visual rendering conversion on the data layout result through the visual rendering layer to generate a visual image.

9. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory stores computer-readable instructions, and when the processor executes the computer-readable instructions, the steps of the image generation method based on depth-aware cyclic attention are implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the image generation method based on depth-aware recurrent attention according to any one of claims 1 to 7.