Image transmission method, image generation method, and storage medium
By conducting semantic analysis and personalized screening and transmission of drone images, combined with generative artificial intelligence to reconstruct images, the problem of low data transmission efficiency in drone patrol images is solved, and efficient and accurate image transmission and personalized reconstruction are achieved.
Patent Information
- Application Number
- CN202510396511.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-18
AI Technical Summary
In the communication environment with huge data volume and limited bandwidth, the data transmission efficiency is low, and the existing technology is difficult to effectively improve.
By semantic analysis of the images collected by the drone, the target description information and the target semantic segmentation map are generated, and image information with high similarity is selected according to the user's keywords for transmission, and the image is reconstructed using generative artificial intelligence to meet the user's personalized needs.
In an environment with limited bandwidth, efficient image data transmission is achieved, the efficiency of image return to the drone inspection is optimized, the user's personalized needs is met, and the accuracy and efficiency of information transmission is improved.
Smart Images

Figure CN120339875A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communications, and in particular, to a method for sending an image, a method for generating an image, and a storage medium. Background Art
[0002] Currently, the wide application of drone technology, especially in scenarios such as infrastructure inspection and environmental monitoring, has become an important means for real-time data collection and analysis. The image acquisition ability of drones enables them to obtain high-resolution images and videos, providing detailed information for subsequent data processing and decision-making. However, with the sharp increase in the amount of image data collected during drone inspections, how to efficiently and quickly transmit this data under limited communication bandwidth has become a technical problem that urgently needs to be solved.
[0003] Traditional communication methods often use data compression or encryption technologies to address the challenge of large amounts of data, but these technologies have obvious limitations in the backhaul of drone images. First, data compression may lead to a decrease in image quality, affecting the accuracy of information and the reliability of decision-making. Second, although encryption can ensure data security, it increases the processing complexity and transmission delay. More importantly, existing semantic communication solutions lack the ability to respond to users' personalized needs when processing drone inspection images, and it is difficult to optimize data transmission according to users' specific needs, thus unable to effectively improve data transmission efficiency in a bandwidth-constrained environment.
[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] The present application provides a method for sending an image, a method for generating an image, and a storage medium, so as to at least solve the technical problem of data transmission efficiency in the backhaul of drone inspection images in a communication environment with a large amount of data and limited bandwidth in the prior art.
[0006] According to one aspect of the present application, a method for sending images is provided, including: collecting N target images, where N is an integer greater than or equal to 1; performing semantic analysis on the N target images to obtain N pieces of first image information, where the first image information includes target description information of the target image and a target semantic segmentation map of the target image, where the target description information is used to characterize the content information in the target image, and the target semantic segmentation map is used to represent the semantic category of each pixel in the target image in the form of an image, and the target description information of each target image is correlated with the target semantic segmentation map of the target image; determining S pieces of target image information according to the target keyword and the N pieces of first image information, where S is an integer greater than or equal to 1 and less than or equal to N, the target keyword is used to characterize the requirements of the target user for the image content, and the S pieces of target image information are the first image information in the N pieces of first image information whose similarity with the target keyword is greater than a preset threshold; sending the S pieces of target image information to the target receiving end.
[0007] Optionally, performing semantic analysis on the N target images to obtain N pieces of first image information includes: determining N pieces of target description information according to the first model and the N target images, where the first model is used to extract image features from the target image and generate text description information associated with the target image based on the image features; determining N target semantic segmentation maps according to the second model and the N target images, where the second model is used to extract the semantic features of each pixel from the target image and assign the semantic features to predefined semantic categories to generate a pixel-by-pixel classified semantic segmentation map; determining N pieces of first image information according to the N pieces of target description information and the N target semantic segmentation maps.
[0008] Optionally, determining S pieces of target image information according to the target keyword and the N pieces of first image information includes: determining the semantic similarity between the target keyword and the target semantic segmentation map of each piece of first image information in the N pieces of first image information; determining S pieces of target image information according to the semantic similarity between the target keyword and the target semantic segmentation map of each piece of first image information.
[0009] Optionally, determining S pieces of target image information according to the semantic similarity between the target keyword and the target semantic segmentation map of each piece of first image information includes: if the semantic similarity between the target semantic segmentation map of the i-th piece of first image information and the target keyword is less than the preset threshold, determining that the i-th piece of first image information is not target image information, where i is an integer greater than or equal to 1 and less than or equal to N; if there are S pieces of first image information in the N pieces of first image information whose semantic similarity with the target keyword is greater than or equal to the preset threshold, determining the S pieces of first image information as the S pieces of target image information.
[0010] According to another aspect of the present application, a method for generating an image is further provided, including: receiving S pieces of target image information, where the S pieces of target image information include: S pieces of target description information and S target semantic segmentation maps, where the S pieces of target description information and the S target semantic segmentation maps are determined by N target images and target keywords of a target user, where the target images are collected by a target sending end, the target description information is used to represent the content information in the target images, the target semantic segmentation maps are used to represent the semantic categories of each pixel in the target images in the form of images, and the target description information of each target image is correlated with the target semantic segmentation map of the target image; generating a reconstructed image according to the S pieces of target image information.
[0011] Optionally, after receiving the S pieces of target image information, the method further includes: performing a preprocessing operation on the target description information, where the preprocessing operation includes: a quantifier removal operation and a control word addition operation, where the quantifier removal operation is used to remove the quantifiers in the target description information, and the control word addition operation is used to add extended information to the target description information; performing a segmentation operation on the target semantic segmentation map according to a preset size, where the segmentation operation is used to segment the size of the target semantic segmentation map to meet the input size requirements of a target neural network.
[0012] Optionally, generating a reconstructed image according to the S pieces of target image information includes: obtaining a user picture, where the user picture is used to reflect the preferences of the target user; inputting the user picture and the S target semantic segmentation maps into a target neural network for feature extraction to obtain target features, where the target neural network fuses the user picture and the S target semantic segmentation maps during the image generation process to guide the target model to generate a reconstructed image that meets preset requirements, where the preset requirements are used to constrain the difference value between the generated reconstructed image and the target image to be less than a preset value; converting the target features into guiding information, where the guiding information includes the preference information of the target user and the structural information and layout information of the target semantic segmentation map, and the guiding information is used to guide the target model to generate a reconstructed image; inputting the S pieces of target description information and the guiding information into the target model to generate a reconstructed image reconstructed based on the target image information, where the target model uses the S pieces of target description information as a reference and generates a reconstructed image according to the guiding information.
[0013] According to another aspect of the present application, there is also provided a transmitting device for images, including: an acquisition unit that acquires N target images, where N is an integer greater than or equal to 1; an analysis unit that performs semantic analysis on the N target images to obtain N pieces of first image information, where the first image information includes target description information of the target images and target semantic segmentation maps of the target images. The target description information is used to represent the content information in the target images, and the target semantic segmentation maps are used to represent the semantic categories of each pixel in the target images in the form of images. Among them, the target description information of each target image is correlated with the target semantic segmentation map of this target image; a determination unit that determines S pieces of target image information according to the target keywords and the N pieces of first image information, where S is an integer greater than or equal to 1 and less than or equal to N. The target keywords are used to represent the requirements of the target user for the image content, and the S pieces of target image information are the first image information among the N pieces of first image information whose similarity to the target keywords is greater than a preset threshold; a transmitting unit that transmits the S pieces of target image information to the target receiving end.
[0014] According to another aspect of the present application, there is also provided a generating device for images, including: a receiving unit that receives S pieces of target image information, where the S pieces of target image information include: S pieces of target description information and S target semantic segmentation maps. The S pieces of target description information and the S target semantic segmentation maps are determined by N target images and the target keywords of the target user. The target images are acquired by the target transmitting end. The target description information is used to represent the content information in the target images, and the target semantic segmentation maps are used to represent the semantic categories of each pixel in the target images in the form of images. Among them, the target description information of each target image is correlated with the target semantic segmentation map of this target image; a generating unit that generates a reconstructed image according to the S pieces of target image information.
[0015] According to another aspect of the present application, there is also provided a computer-readable storage medium. The computer-readable storage medium includes a stored executable program. When the executable program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned image transmitting method or image generating method.
[0016] According to another aspect of the present application, there is also provided an electronic device. The electronic device includes one or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement running the program. Among them, the program is set to execute the above-mentioned image transmitting method or image generating method when running.
[0017] In this application, first, N target images are collected, where N is an integer greater than or equal to 1. Then, semantic analysis is performed on the N target images to obtain N pieces of first image information. The first image information includes the target description information of the target image and the target semantic segmentation map of the target image. The target description information is used to characterize the content information in the target image, and the target semantic segmentation map is used to represent the semantic category of each pixel in the target image in the form of an image. The target description information of each target image is correlated with the target semantic segmentation map of the target image. Then, S pieces of target image information are determined according to the target keyword and the N pieces of first image information, where S is an integer greater than or equal to 1 and less than or equal to N. The target keyword is used to characterize the requirements of the target user for the image content. The S pieces of target image information are the first image information in the N pieces of first image information whose similarity to the target keyword is greater than a preset threshold. Finally, the S pieces of target image information are sent to the target receiving end. That is, through the method of semantic encoding and personalized selection, the purpose of efficient transmission of image data in a communication environment with a large amount of data and limited bandwidth is achieved, thus realizing the technical effect of optimizing the image transmission efficiency of UAV patrol and meeting the personalized needs of users, and further solving the technical problem of the data transmission efficiency of UAV patrol image transmission in a communication environment with a large amount of data and limited bandwidth. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0019] Figure 1 is a flowchart of an optional method for sending images according to an embodiment of the present application Figure 1 ;
[0020] Figure 2 is a flowchart of an optional method for sending images according to an embodiment of the present application Figure 2 ;
[0021] Figure 3 is a flowchart of an optional method for sending images according to an embodiment of the present application Figure 3 ;
[0022] Figure 4 is a flowchart of an optional method for generating images according to an embodiment of the present application;
[0023] Figure 5 is a schematic diagram of an optional method for generating images according to an embodiment of the present application;
[0024] Figure 6Schematic diagram of an optional image sending device according to an embodiment of the present application;
[0025] Figure 7 Schematic diagram of an optional image generating device according to an embodiment of the present application. Detailed implementation manners
[0026] In order to enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0028] It should also be noted that the information and data collected in the present application are information and data authorized by the user or fully authorized by all parties, and the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure and application complies with the relevant laws, regulations and standards of the relevant regions, takes necessary confidentiality measures, does not violate public order and good customs, and provides corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is provided between the present system and relevant users or institutions. Before obtaining relevant information, a request for obtaining information needs to be sent to the aforementioned users or institutions through the interface, and after receiving the consent information fed back by the aforementioned users or institutions, the relevant information is obtained.
[0029] According to an embodiment of the present application, a method embodiment of an image sending method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that here.
[0030] It should be noted that the target sender, such as a drone, is the execution entity of the image sending method in the embodiments of the present application. It can be understood that the image sending method provided in the embodiments of the present application can also be executed by other systems or devices, and the embodiments of the present application do not make specific limitations in this regard.
[0031] Figure 1 is the flow of an optional image sending method according to the embodiments of the present application Figure 1 , such as Figure 1 As shown, the method includes the following steps:
[0032] Step S101, collect N target images.
[0033] In step S101, N is an integer greater than or equal to 1.
[0034] Optionally, the N target images refer to the image dataset collected by the target sender (such as a drone) during the inspection task. These images can be multi-perspective and multi-scene, covering all possible images encountered in the drone inspection task.
[0035] Step S102, perform semantic analysis on the N target images to obtain N pieces of first image information.
[0036] In step S102, the first image information includes the target description information of the target image and the target semantic segmentation map of the target image.
[0037] In step S102, the target description information is used to characterize the content information in the target image, and the target semantic segmentation map is used to represent the semantic category of each pixel in the target image in the form of an image.
[0038] In step S102, the target description information of each target image is correlated with the target semantic segmentation map of that target image.
[0039] Optionally, the target description information: is generated by an image description model (such as ViT-GPT2) on the drone side and is used to summarize the content in the image; the target semantic segmentation map: is generated by a semantic segmentation model (such as OneFormer) on the drone side, classifying each pixel in the image into specific semantic categories, such as vehicles, roads, buildings, etc., for precise positioning and identification of different elements in the image.
[0040] Optionally, the semantic segmentation map is a result of image processing in computer vision. It is obtained from the original image after semantic segmentation processing, where each pixel is classified and labeled as belonging to a specific semantic category. In semantic segmentation, the semantic categories can be "sky", "building", "road", "vehicle", "pedestrian", etc. These categories define different objects or scene elements in the image. The semantic segmentation map is usually presented in the form of a grayscale image or a color image, where different colors represent different semantic categories for easy visual understanding and subsequent processing. For example, in the scenario of drone inspection, the original image may contain a complex environment and multiple objects. After semantic segmentation processing, a semantic segmentation map can be obtained. In this map, each pixel is labeled as an object belonging to a certain category. For example, vehicles on the road are labeled as "car", and buildings are labeled as "building". This enables the system to understand and identify the position and type of specific objects in the image, rather than just treating it as a whole image for processing.
[0041] Optionally, the target sending end performs semantic analysis on each collected target image to obtain the target description information and the target semantic segmentation map corresponding to each image.
[0042] Step S103: Determine S target image information according to the target keyword and N pieces of first image information.
[0043] In step S103, S is an integer greater than or equal to 1 and less than or equal to N. The target keyword is used to characterize the requirements of the target user for the image content. The S pieces of target image information are the first image information among the N pieces of first image information whose similarity with the target keyword is greater than a preset threshold.
[0044] Optionally, the target keyword refers to the keyword of the target user's interest point, such as "building", "car", etc., which is used to accurately express the user's preference for the image content.
[0045] Optionally, the target sending end selects the first image information that meets the requirements of the target user's preference from the N pieces of first image information according to the target keyword of the target user as the target image information.
[0046] Optionally, the target sending end realizes the screening of the user's personalized needs according to the user's keyword, ensuring that only the information highly relevant to the user's preference is further processed and transmitted. By comparing the target keyword with the target description information in the first image information, the target sending end can select the image information that best meets the user's interest. At the same time, using the category information in the semantic segmentation map, the relevance between the image content and the keyword is further confirmed, improving the accuracy of the screening.
[0047] Step S104: Send the S pieces of target image information to the target receiving end.
[0048] Optionally, after the screening is completed, the target sending end only transmits the S image information that is highly relevant to the user's preferences, rather than the original N target images. This not only significantly reduces the amount of transmitted data, reduces the consumption of bandwidth and computing resources, but also improves the efficiency and pertinence of information transmission.
[0049] Optionally, the target receiving end refers to the ground user equipment that communicates with the target sending end and is capable of receiving and processing the semantic information transmitted by the drone.
[0050] Optionally, in an actual scenario, there may be multiple users in an environment with limited communication. Each user has a certain need for the drone inspection image backhaul, and according to their own interests, the required images are also different, more personalized and preference-based. At the same time, the user's needs will also change dynamically over time. Therefore, it is also necessary to meet the dynamic needs of users when collecting images. Figure 2 It is the flow of an optional image sending method according to an embodiment of the present application Figure 2 , as Figure 2 shown. First, the user outputs a subscription keyword on the user's ground equipment. Before the drone inspection, the drone will collect the subscription keyword regularly to obtain the user's preference information for subsequent information screening during image backhaul. After the user outputs the subscription keyword, the drone platform collects the user's preference information (subscription keyword), especially the keywords subscribed by the user in different scenarios (such as interest topics, content types, etc.). Then, when the drone conducts inspections, it will perform tasks such as image acquisition and semantic analysis. After that, the drone will screen the backhaul images according to the collected subscription keywords of the user and transmit the images that meet the user's preferences. Finally, the user obtains the preference content encoded with semantics on the user side (user's ground equipment).
[0051] As can be seen from the content of steps S101 to S104, in this application, first, N target images are collected, where N is an integer greater than or equal to 1. Then, semantic analysis is performed on the N target images to obtain N pieces of first image information. The first image information includes the target description information of the target image and the target semantic segmentation map of the target image. The target description information is used to represent the content information in the target image, and the target semantic segmentation map is used to represent the semantic category of each pixel in the target image in the form of an image. The target description information of each target image is correlated with the target semantic segmentation map of that target image. Then, S pieces of target image information are determined according to the target keyword and the N pieces of first image information, where S is an integer greater than or equal to 1 and less than or equal to N. The target keyword is used to represent the requirements of the target user for the image content. The S pieces of target image information are the first image information in the N pieces of first image information whose similarity to the target keyword is greater than a preset threshold. Finally, the S pieces of target image information are sent to the target receiving end. That is, through the method of semantic encoding and personalized selection, the purpose of efficient transmission of image data in a communication environment with a large amount of data and limited bandwidth is achieved, thereby realizing the technical effects of optimizing the image transmission efficiency of drone patrol and meeting the personalized needs of users, and further solving the technical problem of the data transmission efficiency of drone patrol image transmission in a communication environment with a large amount of data and limited bandwidth.
[0052] In an alternative embodiment, the target sending end first determines N pieces of target description information according to the first model and the N target images. The first model is used to extract image features from the target images and generate text description information associated with the target images based on the image features. Then, N target semantic segmentation maps are determined according to the second model and the N target images. The second model is used to extract the semantic features of each pixel from the target images and assign the semantic features to predefined semantic categories to generate a pixel-by-pixel classified semantic segmentation map. Finally, N pieces of first image information are determined according to the N pieces of target description information and the N target semantic segmentation maps.
[0053] Optionally, the first model refers to an image description model, such as vit-gpt2-image-captioning (Vision Transformer with Generative Pre-Training for Image Captioning), which is used to extract image features from the target images and generate text information describing the content of the target images based on these features.
[0054] Optionally, the second model refers to a semantic segmentation model, such as the oneformer model based on transformer (a neural network architecture), which is a model for pixel-level prediction tasks (such as semantic segmentation, instance segmentation, etc.). It performs semantic segmentation on the collected image, converts it into a semantic segmentation map, and assigns each pixel in the image to a specific category, thereby classifying the entire image pixel by pixel and generating a pixel-level classification map.
[0055] Optionally, after the image acquisition device carried by the target sender captures N target images in the inspection area, each target image is input into the first model to obtain the text description information corresponding to the target image. Among them, these text information contains the key content information of the target image, which is used to assist the subsequent image screening and personalized image generation process. Through the first model, the target sender can convert the image into a concise and understandable text description for subsequent processing; at the same time, each target image is input into the second model, and the second model performs semantic segmentation on each input target image, identifies and marks the semantic category to which each pixel belongs, and obtains the semantic segmentation map corresponding to the target image. Among them, the generated semantic segmentation map not only contains the object category but also retains the position information of the object in the image, which provides the basis for the structure and layout of the subsequent image reconstruction. By generating the target semantic segmentation map, the drone can provide a detailed image content structure, which is crucial for semantic-based image transmission and reconstruction.
[0056] Optionally, the final target sender pairs the N target description information generated in the first step and the N target semantic segmentation maps generated in the second step to form N first image information. Each first image information contains the semantic text description of the target image and the visual representation of the semantic segmentation, preparing materials for the subsequent user preference analysis and personalized transmission of image information.
[0057] Optionally, there are multiple drone agents in the actual scenario. After collecting user preference information, they will patrol multiple specified scenarios and collect images through the equipped image collection device. For the collected images, semantic information analysis will be carried out, and the specific implementation is divided into an image-based text description model and a semantic segmentation model. Figure 3 It is the flow of an optional image sending method according to an embodiment of the present application Figure 3 such as Figure 3As shown, after the drone captures an image, it processes the image based on a text description model and a semantic segmentation model. Through the text description model, the text description content of the image can be obtained, such as "Car driving on the highway". Through the semantic segmentation model, the semantic content of the image can be conveniently extracted. By dividing different regions in the image into parts with specific semantics, such as buildings, roads, pedestrians, etc., only the important regional information related to the target task or user needs is retained, and finally a semantic segmentation map is obtained. Then, it is filtered according to the user's preferences. Finally, the filtered image information (the text description content and the semantic segmentation map corresponding to each filtered image) is sent back to the user.
[0058] As can be seen from the above, by introducing the first model (image description model) and the second model (semantic segmentation model), the target sending end can automatically extract and generate semantic text descriptions and detailed semantic segmentation maps from the target image, and then integrate them to form N first image information. The combined use of these two models enables the target sending end not only to quickly generate a general description of the image, but also to carefully analyze the semantic categories of each pixel in the image, providing a solid foundation for subsequent personalized image screening, transmission, and reconstruction. And through the above implementation methods, efficient compression and semantic enhancement of image information are achieved, which can significantly reduce the communication load and improve the transmission efficiency in the UAV inspection scenario with a large amount of data and limited bandwidth. At the same time, based on the generated semantic information, users can express their points of interest and needs more precisely, and the target sending end can also respond more intelligently to the user's preferences and generate high-quality reconstructed images as expected by the user, thus significantly improving the user experience and the effectiveness of communication.
[0059] In an alternative embodiment, the target sending end first determines the semantic similarity between the target keyword and the target semantic segmentation map of each of the N first image information, and then determines S target image information according to the semantic similarity between the target keyword and the target semantic segmentation map of each first image information.
[0060] Optionally, for each of the N first image information, the target sending end compares the target semantic segmentation map with the target keyword and calculates the semantic similarity between the two. This usually involves matching the keyword with the semantic categories appearing in the semantic segmentation map, and similarity can be quantified using algorithms such as cosine similarity or specific image matching algorithms (such as deep learning-based feature matching). The calculated semantic similarity reflects the degree of content association between the keyword and the content in each target semantic segmentation map, providing a basis for subsequent image information screening.
[0061] Optionally, the target sender filters out the set of image information that is most relevant to the target keyword from the N pieces of first image information according to the calculated similarity and a preset similarity threshold.
[0062] As can be seen from the above, first, by calculating the semantic similarity between the target keyword and the target semantic segmentation map, the target sender can accurately identify which image information is most relevant to the user's needs. Second, the S pieces of target image information selected by similarity ensure that only the content that the user truly cares about is transmitted, greatly reducing bandwidth occupancy and transmission latency. At the same time, it also avoids the user receiving a large amount of irrelevant information and reduces the complexity of subsequent processing. That is, it realizes the screening of image information based on user preferences, improves the efficiency of image backhaul and the user experience. It not only reduces the amount of data transmitted, but also ensures the high relevance of the transmitted content, making information transmission more accurate and efficient. Especially in complex scenarios with multiple users and multiple demands, it can effectively meet the personalized image needs of different users, reduce the overall load of the communication system, and is a key step in improving bandwidth utilization efficiency and user experience in the scenario of drone inspection image backhaul. Compared with traditional full-volume image transmission, this solution significantly reduces the transmission of non-critical information, thereby reducing the cost and time of data transmission, while maintaining the integrity and accuracy of image information and meeting the requirements of real-time and efficient information transmission.
[0063] In an optional embodiment, if the semantic similarity between the target semantic segmentation map of the i-th piece of first image information and the target keyword is less than the preset threshold, the target sender determines that the i-th piece of first image information is not target image information, where i is an integer greater than or equal to 1 and less than or equal to N. If there are S pieces of first image information among the N pieces of first image information whose semantic similarities with the target keyword are all greater than or equal to the preset threshold, the target sender determines the S pieces of first image information as S pieces of target image information.
[0064] Optionally, if the calculated semantic similarity between the target semantic segmentation map of the i-th piece of first image information and the target keyword is less than the preset threshold, the target sender automatically excludes this image information and does not regard it as target image information, thus avoiding unnecessary subsequent transmission and processing.
[0065] Optionally, the S pieces of target image information: a set of image information filtered out from the N pieces of first image information whose semantic similarity with the target keyword is higher than the preset threshold. The size of S may be smaller than N, but at least 1.
[0066] Optionally, the S pieces of first image information constitute the target image information that is ultimately to be transmitted to the user. They meet the specific requirements of the user for the image content and are high-quality and highly relevant image information.
[0067] As can be seen from the above, the target sending end dynamically filters out image information that does not match the user's preferences by calculating the semantic similarity between the target semantic segmentation map and the target keywords, reducing unnecessary data transmission and consumption of communication bandwidth and computing resources. By only transmitting S pieces of target image information, the target sending end can significantly improve the efficiency of information transmission, ensuring that the image information received by the user not only meets the requirements in terms of quantity but also highly conforms to their preferences in terms of content, improving user satisfaction and the overall performance of the system. And in the scenario of drone inspection, this screening mechanism based on semantic similarity is of great significance for optimizing the image backhaul process. It not only solves the transmission bottleneck problem caused by a large amount of data but also enables the system to flexibly adjust the backhaul strategy according to the dynamically changing needs of the user, reflecting the innovative advantages of this embodiment in improving communication efficiency and user experience. Through precise semantic matching and intelligent image information screening, the drone inspection system can more effectively serve the personalized needs of multiple users, reducing information overload and improving the utilization rate of the entire communication link.
[0068] According to an embodiment of the present application, a method embodiment of an image generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0069] It should be noted that the target receiving end, as the execution subject of the image sending method of this embodiment of the present application, such as a ground user device. It can be understood that the image generation method provided by this embodiment of the present application can also be executed by other systems or devices as the execution subject, and this embodiment of the present application does not make specific limitations on this.
[0070] Figure 4 is a flowchart of an optional image generation method according to an embodiment of the present application, as Figure 4 shown, the method includes the following steps:
[0071] Step S401, receive S pieces of target image information.
[0072] In step S401, the S pieces of target image information include: S pieces of target description information and S target semantic segmentation maps.
[0073] In step S401, the S pieces of target description information and the S target semantic segmentation maps are determined by N target images and the target keywords of the target user.
[0074] In step S401, the target image is collected by the target sender. The target description information is used to characterize the content information in the target image, and the target semantic segmentation map is used to represent the semantic category of each pixel in the target image in the form of an image.
[0075] In step S401, the target description information of each target image is correlated with the target semantic segmentation map of that target image.
[0076] Optionally, the S target image information: These information are highly relevant to the keywords subscribed by the user and are S target description information and S target semantic segmentation maps determined according to the user's keywords screening and semantic similarity calculation among the N target images collected by the target sender.
[0077] Optionally, there is a close semantic connection between the target description information and the target semantic segmentation map. The target description information is a textual description of the content of the target semantic segmentation map, while the target semantic segmentation map is a visual expression of the target description information. The two together constitute a comprehensive understanding of the content of the target image.
[0078] Optionally, in this step, the target receiver receives the S target image information transmitted by the target sender. These information are the image data that best match the user's preferences after semantic analysis and keyword matching. The combination of the target description information and the target semantic segmentation map provides rich semantic information for the target receiver and lays a foundation for the subsequent generation of the reconstructed image.
[0079] Step S402, generate a reconstructed image according to the S target image information.
[0080] Optionally, the target receiver can use generative artificial intelligence technology to generate a high-quality image that meets the user's expectations based on the received target description information and target semantic segmentation map.
[0081] In an optional embodiment, the target receiver first performs a preprocessing operation on the target description information. Among them, the preprocessing operation includes: a quantifier removal operation and a control word addition operation. The quantifier removal operation is used to remove the quantifiers in the target description information, and the control word addition operation is used to add extended information to the target description information. Then, the target semantic segmentation map is segmented according to a preset size. The segmentation operation is used to segment the size of the target semantic segmentation map to meet the input size requirements of the target neural network.
[0082] Optionally, the quantifier removal operation: refers to removing the words indicating quantity (such as "one", "multiple", etc.) from the target description information to reduce redundant information in the text description; the control word addition operation: adding specific control words to the target description information, which can be used to guide the direction of image generation, provide more specific detailed descriptions, and enhance the controllability and precision of image generation.
[0083] Optionally, after receiving S target description information, the target receiving end performs the quantifier removal operation on each target description information, deleting unnecessary quantity information in the description to make the text more concise. Then, for each processed target description information, the target receiving end performs the control word addition operation, adding words such as "clear" and "detail-rich" that are helpful for controlling the image generation process according to user preferences to improve the quality of image reconstruction.
[0084] Optionally, after receiving S target semantic segmentation maps, the target receiving end checks whether the size of each semantic segmentation map meets the preset size requirements. For the semantic segmentation maps with mismatched sizes, the target receiving end performs segmentation operations such as cropping, scaling, or padding to adjust them to the preset size. After the segmentation operation, the sizes of the target semantic segmentation maps are standardized and can be used as conditional inputs for subsequent neural networks, providing detailed guidance on the structure and layout for image reconstruction.
[0085] As can be seen from the above, through preprocessing and segmentation operations, the target receiving end ensures that it can more effectively utilize the received target description information and target semantic segmentation maps for image reconstruction. By removing quantifiers and adding control words, the preprocessing operation optimizes the text description, making it more refined and instructive, which helps to generate reconstructed images closer to user expectations. At the same time, the segmentation operation adjusts the target semantic segmentation maps to the preset size that meets the model input, ensuring the smooth progress of the image generation process and avoiding model processing problems caused by size mismatches. Generally speaking, these preprocessing and segmentation operations work together to improve the accuracy and efficiency of image reconstruction, reduce unnecessary data processing steps, and ensure the accurate transmission and realization of user personalized preferences. In the scenarios of drone image backhaul and semantic communication, this technical solution can significantly improve user satisfaction with the reconstructed images, optimize the utilization of communication resources, and reduce the complexity and computational cost of image reconstruction. Through refined image information processing, this embodiment not only optimizes the image transmission process but also improves the quality of image reconstruction based on generative artificial intelligence, providing users with more accurate and efficient information services.
[0086] In an alternative embodiment, the target receiving end first obtains a user picture, where the user picture is used to reflect the preferences of the target user. Then, the user picture and S target semantic segmentation maps are input into a target neural network for feature extraction to obtain target features. The target neural network guides the target model to generate a reconstructed image that meets the preset requirements by fusing the user picture and S target semantic segmentation maps during the image generation process. The preset requirements are used to constrain the difference value between the generated reconstructed image and the target image to be less than a preset value. Then, the target features are converted into guidance information, where the guidance information includes the preference information of the target user, as well as the structural information and layout information of the target semantic segmentation maps. The guidance information is used to guide the target model to generate a reconstructed image. Finally, the S target description information and the guidance information are input into the target model to generate a reconstructed image based on the target image information. The target model generates the reconstructed image based on the guidance information with the S target description information as a reference.
[0087] Optionally, the target receiving end collects pictures provided by the user. These pictures can be images that the user has browsed historically, preference images for specific scenarios, or any images that can express the user's visual preferences and content requirements. The user picture is an important reference for generating the reconstructed image, helping the system understand the user's special requirements for the image, such as style, color, or the type of scene to be focused on.
[0088] Optionally, the target neural network is ControlNet (Controlled Network). ControlNet is a neural network structure used to enhance the control ability of an image generation model, especially when combined with a diffusion model. It makes the image generation process more controllable by introducing additional conditions (such as edge detection maps, depth maps, human poses, etc.) during the image generation process. This model architecture can increase the precise control of specific elements in the generated image while maintaining the powerful capabilities of the original diffusion model. The design concept of ControlNet is to use the deep encoding layer of a large-scale pre-trained model as a powerful foundation and add "zero convolutions" (zero-initialized convolutional layers) to gradually increase parameters, ensuring that no harmful noise is introduced during the fine-tuning process. In this way, ControlNet can handle various types of conditional control and achieve robust training effects on both small-scale and large-scale datasets. In the application of this embodiment, ControlNet allows users to precisely control the components of the generated image by providing control images (semantic segmentation maps and background knowledge reference maps).
[0089] Optionally, the target model is the Stable Diffusion model. Stable Diffusion is a model based on a probabilistic diffusion process. It generates images by gradually denoising and can perform excellently in high-resolution image generation. Its working principle is to gradually remove the noise and reversely restore the structure and details in the image, finally generating a clear image. The key to Stable Diffusion is to learn how to effectively denoise at each step, so that the finally generated image has high fidelity in both details and overall structure. Compared with other generative models, Stable Diffusion has strong stability and flexibility, and can handle complex image generation tasks, being applicable to various image generation and reconstruction scenarios. The introduction of Stable Diffusion makes image generation more flexible, especially suitable for tasks that require higher image resolution.
[0090] Optionally, the target receiver takes the user picture and S target semantic segmentation maps as inputs and feeds them into the target neural network. The target neural network performs feature extraction. According to the design of ControlNet, the network first processes the user picture to extract its style and content features; at the same time, it also processes the target semantic segmentation maps to understand the content structure of the images. In the feature fusion stage, the neural network integrates the features of the user picture with the structural and layout information of the target semantic segmentation maps to obtain target features that can simultaneously reflect the user's preferences and the semantic content of the target images.
[0091] Optionally, the target neural network converts the extracted target features into guidance information. This guidance information not only contains the style and preference information conveyed by the user picture, but also contains the detailed information of the key structural layouts in the target semantic segmentation maps. The construction of the guidance information ensures that when the target model generates a reconstructed image, it can comprehensively consider the personalized needs of the user and the semantic structure of the original image, and generate a new image that not only meets the user's preferences but also maintains the semantic integrity of the original image.
[0092] Optionally, the target receiver takes S target description information as text inputs and the guidance information as conditional inputs and feeds them into the target model. The target model generates a reconstructed image that is semantically similar to the original target image and whose style and details meet the user's expectations based on the text description of the S target description information and the user preferences and semantic structure information contained in the guidance information. The generated reconstructed image not only retains the key information of the original image, but also makes personalized adjustments according to the specific preferences of the user, achieving a dual optimization of image content and style.
[0093] Optionally, Figure 5 is a schematic diagram of an optional method for generating an image according to an embodiment of the present application, asFigure 5 As shown, the text description information of each image in the image information transmitted back from the drone is used as a prompt, the semantic segmentation map is used as a reference map, and the user-side background knowledge (user pictures) is considered as a reference. The ControlNet and Stable Diffusion models are used in combination to achieve fine control and enhancement in image generation and realize image reconstruction. Specifically, first, the obtained text description information is adjusted, such as removing quantifiers and adding control words, and the semantic segmentation map is cropped to obtain an image that matches the input size of the user-side model. Then, the semantic segmentation map and the user-side background knowledge are input into ControlNet for processing, and the output of ControlNet and the text description information are input into StableDiffusion for image generation. With the help of the user-side computing power and the pre-trained generation model, a reconstructed image that meets the user's expectations is generated, completing semantic communication.
[0094] Optionally, the image generation model based on generative artificial intelligence in this embodiment has strong image generation capabilities and strong information generation capabilities. It can fill in the missing parts according to the context when transmitting data, and has higher robustness and adaptability. When the data transmission is incomplete or the signal quality is poor, our solution can generate reasonable content based on the existing semantics for supplementation. Moreover, the fine control of image generation by ControlNet is utilized (mainly in two parts: one part is using the semantic segmentation map as an image reference, and the other part is using user background knowledge), making the reconstructed image reach a similar style and effect. At the same time, based on the pre-trained base model, it can dynamically generate content based on different instructions and seamlessly switch in multiple tasks and scenarios without the need for retraining.
[0095] As can be seen from the above, by introducing user pictures and the target neural network, image reconstruction based on user preferences is achieved. The acquisition of user pictures enables the target receiving end to understand and capture the user's personalized needs, and the generated guidance information by fusing the user picture features and the structural information of the target semantic segmentation map through the target neural network ensures that the reconstructed image meets both content accuracy and style matching. During the process of generating the reconstructed image, the target model is comprehensively guided according to the user preference information and the semantic information of the target image. The generated image not only conveys the core semantics of the original image but also reflects the user's personalized requirements, such as specific visual styles, color preferences, or scene elements of concern. This personalized image generation based on user preferences significantly improves the user experience. Especially in the drone inspection scenario, it can provide accurate and personalized content services according to the dynamic needs of different users, reduce unnecessary data transmission and processing, and optimize the use of communication resources.
[0096] In an alternative embodiment, in the scenario of drone inspection, the drone device carries an image collection device to collect information in the environment and transmit it back to the user for image transmission. However, in the scenario with limited communication, it is impossible to rely on the traditional communication transmission scheme for information transmission. This embodiment adopts a semantic encoding and decoding scheme based on generative artificial intelligence, which has higher flexibility and intelligence. Briefly speaking, this semantic encoding and decoding scheme uses a semantic segmentation model as the encoder to convert the image into a semantic segmentation map, thereby achieving better semantic extraction and compression. On the semantic decoding side, this embodiment utilizes generative artificial intelligence, that is, an image reconstruction scheme combining the ControlNet and Stable Diffusion models, which can achieve better semantic reconstruction effects based on the image semantic segmentation map. At the same time, optimization is also performed based on user preferences when the image is transmitted back. The drone will collect the subscription interest keywords on the user side. For example, some users may be interested in vehicles, while some are interested in buildings. After that, when the drone conducts inspections and collects corresponding images, they will be transmitted separately according to the different needs of users instead of being all transmitted back, significantly reducing the communication load and facilitating the use on the user side. After the drone cluster runs for a certain period of time, it will also re-collect the subscription words of users to adapt to the dynamically changing needs of users in real time.
[0097] The embodiment of the present application also provides an image sending device. It should be noted that the image sending device in the embodiment of the present application can be used to execute the image sending method provided in the embodiment of the present application. The following introduces the image sending device provided in the embodiment of the present application.
[0098] According to the embodiment of the present application, there is also provided a device for implementing the above-mentioned image sending method. Figure 6 It is a schematic diagram of an alternative image sending device according to the embodiment of the present application, as Figure 6 shown. The device includes: a collection unit 601, an analysis unit 602, a determination unit 603, and a sending unit 604.
[0099] Optionally, an acquisition unit 601 is configured to acquire N target images, where N is an integer greater than or equal to 1; an analysis unit 602 is configured to perform semantic analysis on the N target images to obtain N pieces of first image information, where the first image information includes target description information of the target image and a target semantic segmentation map of the target image. The target description information is used to characterize the content information in the target image, and the target semantic segmentation map is used to represent the semantic category of each pixel in the target image in the form of an image. The target description information of each target image is correlated with the target semantic segmentation map of the target image; a determination unit 603 is configured to determine S pieces of target image information according to a target keyword and the N pieces of first image information, where S is an integer greater than or equal to 1 and less than or equal to N, the target keyword is used to characterize the requirements of the target user for the image content, and the S pieces of target image information are the first image information in the N pieces of first image information whose similarity to the target keyword is greater than a preset threshold; a sending unit 604 is configured to send the S pieces of target image information to a target receiving end.
[0100] Optionally, the analysis unit 602 includes: a first determination subunit, a second determination subunit, and a third determination subunit. Among them, the first determination subunit is configured to determine N pieces of target description information according to a first model and the N target images, where the first model is used to extract image features from the target images and generate text description information associated with the target images based on the image features; the second determination subunit is configured to determine N target semantic segmentation maps according to a second model and the N target images, where the second model is used to extract semantic features of each pixel from the target images and assign the semantic features to predefined semantic categories to generate a pixel-by-pixel classified semantic segmentation map; the third determination subunit is configured to determine N pieces of first image information according to the N pieces of target description information and the N target semantic segmentation maps.
[0101] Optionally, the determination unit 603 includes: a fourth determination subunit and a fifth determination subunit. Among them, the fourth determination subunit is configured to determine the semantic similarity between the target keyword and the target semantic segmentation map of each piece of first image information in the N pieces of first image information; the fifth determination subunit is configured to determine S pieces of target image information according to the semantic similarity between the target keyword and the target semantic segmentation map of each piece of first image information.
[0102] Optionally, the fifth determination subunit includes: a first determination module and a second determination module. The first determination module is configured to determine that the i-th first image information is not the target image information if the semantic similarity between the target semantic segmentation map of the i-th first image information and the target keyword is less than a preset threshold, where i is an integer greater than or equal to 1 and less than or equal to N; the second determination module is configured to determine that S first image information among the N first image information are S target image information if the semantic similarity between the S first image information and the target keyword is greater than or equal to the preset threshold.
[0103] An embodiment of the present application also provides an image generation device. It should be noted that the image generation device in the embodiment of the present application can be used to execute the image generation method provided in the embodiment of the present application. The following introduces the image generation device provided in the embodiment of the present application.
[0104] According to an embodiment of the present application, there is also provided a device for implementing the above-mentioned image generation method. Figure 7 is a schematic diagram of an optional image generation device according to an embodiment of the present application, as Figure 7 shown, the device includes: a receiving unit 701 and a generating unit 702.
[0105] Optionally, the receiving unit 701 is configured to receive S target image information, where the S target image information includes: S target description information and S target semantic segmentation maps, where the S target description information and the S target semantic segmentation maps are determined by N target images and the target keyword of the target user, the target image is collected by the target sending end, the target description information is used to characterize the content information in the target image, the target semantic segmentation map is used to represent the semantic category of each pixel in the target image in the form of an image, and the target description information of each target image is correlated with the target semantic segmentation map of the target image; the generating unit 702 is configured to generate a reconstructed image according to the S target image information.
[0106] Optionally, the image generation device further includes: a first processing unit and a second processing unit. The first processing unit is configured to perform a preprocessing operation on the target description information, where the preprocessing operation includes: a quantifier removal operation and an additional control word operation, the quantifier removal operation is used to remove the quantifier in the target description information, and the additional control word operation is used to add extended information to the target description information; the second processing unit is configured to perform a segmentation operation on the target semantic segmentation map according to a preset size, where the segmentation operation is used to divide the size of the target semantic segmentation map into a size that meets the input size requirements of the target neural network.
[0107] Optionally, the generation unit 702 includes: an acquisition subunit, an extraction subunit, a conversion subunit, and a generation subunit. Among them, the acquisition subunit is used to acquire a user picture, where the user picture is used to reflect the preferences of the target user; the extraction subunit is used to input the user picture and S target semantic segmentation maps into a target neural network for feature extraction to obtain target features, where the target neural network fuses the user picture and S target semantic segmentation maps during the image generation process to guide the target model to generate a reconstructed image that meets the preset requirements, and the preset requirements are used to constrain the difference value between the generated reconstructed image and the target image to be less than a preset value; the conversion subunit is used to convert the target features into guidance information, where the guidance information includes the preference information of the target user and the structural information and layout information of the target semantic segmentation maps, and the guidance information is used to guide the target model to generate a reconstructed image; the generation subunit is used to input the S target description information and the guidance information into the target model to generate a reconstructed image based on the target image information, where the target model uses the S target description information as a reference and generates a reconstructed image according to the guidance information.
[0108] According to another aspect of the present application, there is also provided a computer-readable storage medium, and the computer-readable storage medium includes a stored executable program. Among them, when the executable program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned image sending method or image generation method.
[0109] According to another aspect of the present application, there is also provided an electronic device, and the electronic device includes one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, it enables the one or more processors to implement a program for running, where the program is set to execute the above-mentioned image sending method or image generation method when running.
[0110] The serial numbers of the above-mentioned embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0111] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0112] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0113] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0114] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0115] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks, or optical discs that can store program codes.
[0116] The above is only the preferred embodiment of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A method for transmitting an image, characterized in that, Including: Collecting N target images, where N is an integer greater than or equal to 1; Performing semantic analysis on the N target images to obtain N pieces of first image information, where the first image information includes the target description information of the target image and the target semantic segmentation map of the target image. The target description information is used to represent the content information in the target image, and the target semantic segmentation map is used to represent the semantic category of each pixel in the target image in the form of an image. The target description information of each target image is correlated with the target semantic segmentation map of this target image; Determining S pieces of target image information according to the target keyword and the N pieces of first image information, where S is an integer greater than or equal to 1 and less than or equal to N. The target keyword is used to represent the requirements of the target user for the image content, and the S pieces of target image information are the first image information in the N pieces of first image information whose similarity with the target keyword is greater than a preset threshold; Sending the S pieces of target image information to the target receiving end.
2. The method for sending an image according to claim 1, wherein Performing semantic analysis on the N target images to obtain N pieces of first image information, including: Determining N pieces of target description information according to the first model and the N target images, where the first model extracts image features from the target images and generates text description information associated with the target images based on the image features; Determining N target semantic segmentation maps according to the second model and the N target images, where the second model is used to extract the semantic features of each pixel from the target images and assign the semantic features to predefined semantic categories to generate a pixel-by-pixel classified semantic segmentation map; Determining N pieces of first image information according to the N pieces of target description information and the N target semantic segmentation maps.
3. The method for sending an image according to claim 1, wherein Determining S pieces of target image information according to the target keyword and the N pieces of first image information, including: Determining the semantic similarity between the target keyword and the target semantic segmentation map of each first image information in the N pieces of first image information; Determining S pieces of target image information according to the semantic similarity between the target keyword and the target semantic segmentation map of each first image information.
4. The method for sending an image according to claim 1, characterized in that Determining S pieces of target image information according to the semantic similarity between the target keyword and the target semantic segmentation map of each first image information, including: If the semantic similarity between the target semantic segmentation map of the i-th first image information and the target keyword is less than the preset threshold, determining that the i-th first image information is not the target image information, where i is an integer greater than or equal to 1 and less than or equal to N; If there are S pieces of first image information in the N pieces of first image information whose semantic similarity with the target keyword is greater than or equal to the preset threshold, determining that the S pieces of first image information are the S pieces of target image information.
5. A method for generating an image, characterized in that, Including: Receive S target image information, where the S target image information includes: S target description information and S target semantic segmentation maps, where the S target description information and the S target semantic segmentation maps are determined by N target images and the target keywords of the target user. The target images are collected by the target sending end. The target description information is used to represent the content information in the target image, and the target semantic segmentation map is used to represent the semantic category of each pixel in the target image in the form of an image. The target description information of each target image is correlated with the target semantic segmentation map of this target image; Generate a reconstructed image according to the S target image information.
6. The method for generating an image according to claim 5, wherein After receiving the S target image information, the method further includes: Perform a preprocessing operation on the target description information, where the preprocessing operation includes: a quantifier removal operation and an additional control word operation. The quantifier removal operation is used to remove the quantifiers in the target description information, and the additional control word operation is used to add extended information to the target description information; Perform a segmentation operation on the target semantic segmentation map according to a preset size, where the segmentation operation is used to segment the size of the target semantic segmentation map to meet the input size requirements of the target neural network.
7. The method for generating an image according to claim 6, wherein Generating a reconstructed image according to the S target image information includes: Obtain a user picture, where the user picture is used to reflect the preferences of the target user; Input the user picture and the S target semantic segmentation maps into the target neural network for feature extraction to obtain target features. The target neural network guides the target model to generate a reconstructed image that meets the preset requirements by fusing the user picture and the S target semantic segmentation maps during the image generation process. The preset requirements are used to constrain that the difference value between the generated reconstructed image and the target image is less than a preset value; Convert the target features into guidance information, where the guidance information includes the preference information of the target user and the structural information and layout information of the target semantic segmentation map. The guidance information is used to guide the target model to generate the reconstructed image; Input the S target description information and the guidance information into the target model to generate a reconstructed image reconstructed based on the target image information, where the target model generates the reconstructed image with reference to the S target description information and based on the guidance information.
8. An image sending device, characterized in that, Includes: A collection unit that collects N target images, where N is an integer greater than or equal to 1; An analysis unit that performs semantic analysis on the N target images to obtain N first image information, where the first image information includes the target description information of the target image and the target semantic segmentation map of the target image. The target description information is used to represent the content information in the target image, and the target semantic segmentation map is used to represent the semantic category of each pixel in the target image in the form of an image. The target description information of each target image is correlated with the target semantic segmentation map of this target image; A determination unit determines S pieces of target image information according to a target keyword and N pieces of the first image information, where S is an integer greater than or equal to 1 and less than or equal to N, the target keyword is used to characterize the requirements of a target user for image content, and the S pieces of target image information are the first image information among the N pieces of the first image information whose similarity with the target keyword is greater than a preset threshold; A sending unit sends the S pieces of target image information to a target receiving end.
9. The device according to claim 8, characterized in that, The analysis unit includes: A first determination subunit is configured to determine N pieces of target description information according to a first model and N pieces of the target images, where the first model extracts image features from the target images and generates text description information associated with the target images based on the image features; A second determination subunit is configured to determine N pieces of target semantic segmentation maps according to a second model and N pieces of the target images, where the second model extracts semantic features of each pixel from the target images and assigns the semantic features to predefined semantic categories to generate a pixel-by-pixel classified semantic segmentation map; A third determination subunit is configured to determine N pieces of the first image information according to the N pieces of target description information and the N pieces of target semantic segmentation maps.
10. The device according to claim 8, characterized in that, The determination unit includes: A fourth determination subunit is configured to determine the semantic similarity between the target keyword and the target semantic segmentation map of each first image information among the N pieces of the first image information; A fifth determination subunit is configured to determine the S pieces of target image information according to the semantic similarity between the target keyword and the target semantic segmentation map of each first image information.
11. The device according to claim 10, characterized in that, The fifth determination subunit includes: A first determination module is configured to determine that the i-th first image information is not the target image information if the semantic similarity between the target semantic segmentation map of the i-th first image information and the target keyword is less than the preset threshold, where i is an integer greater than or equal to 1 and less than or equal to N; A second determination module is configured to determine that the S pieces of first image information are the S pieces of target image information if there are S pieces of first image information among the N pieces of first image information whose semantic similarity with the target keyword is greater than or equal to the preset threshold.
12. An image generation device, characterized in that, It includes: A receiving unit receives S pieces of target image information, where the S pieces of target image information include: S pieces of target description information and S pieces of target semantic segmentation maps, where the S pieces of target description information and the S pieces of target semantic segmentation maps are determined by N pieces of target images and a target keyword of a target user, where the target images are collected by a target sending end, the target description information is used to characterize the content information in the target images, the target semantic segmentation map is used to represent the semantic category of each pixel in the target images in the form of an image, and the target description information of each target image is correlated with the target semantic segmentation map of the target image; A generating unit generates a reconstructed image according to the S pieces of target image information.
13. The device according to claim 12, characterized in that, The device further includes: A first processing unit for performing preprocessing operations on the target description information, where the preprocessing operations include: a quantifier removal operation and a control word addition operation. The quantifier removal operation is used to remove quantifiers in the target description information, and the control word addition operation is used to add extended information to the target description information; A second processing unit for performing a segmentation operation on the target semantic segmentation map according to a preset size, where the segmentation operation is used to segment the size of the target semantic segmentation map to meet the input size requirements of the target neural network.
14. The device according to claim 12, characterized in that, The generating unit includes: An obtaining subunit for obtaining a user picture, where the user picture is used to reflect the preferences of the target user; An extracting subunit for inputting the user picture and S target semantic segmentation maps into the target neural network for feature extraction to obtain target features. The target neural network guides the target model to generate a reconstructed image that meets preset requirements by fusing the user picture and S target semantic segmentation maps during the image generation process. The preset requirements are used to constrain the difference value between the generated reconstructed image and the target image to be less than a preset value; A converting subunit for converting the target features into guidance information, where the guidance information includes the preference information of the target user and the structural information and layout information of the target semantic segmentation map, and the guidance information is used to guide the target model to generate the reconstructed image; A generating subunit for inputting S target description information and the guidance information into the target model to generate a reconstructed image based on the target image information. The target model generates the reconstructed image with reference to S target description information and according to the guidance information.
15. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When the computer program runs, the device where the computer-readable storage medium is located executes the image sending method according to any one of claims 1 to 4 or the image generating method according to any one of claims 5 to 7.
16. An electronic device, characterized in that, The electronic device includes one or more processors; a storage device for storing one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to implement a program for running, where the program is set to execute the image sending method according to any one of claims 1 to 4 or the image generating method according to any one of claims 5 to 7 when running.
Citation Information
Patent Citations
System and method for controlling video information filtering based on semantic content
CN103258050A
Image detection method and device, electronic equipment and computer readable storage medium
CN116091982A
Image cognition semantic communication system and method based on multi-modal knowledge graph
CN118260432A
Method for generating personal knowledge graph and user device using the same
KR1020260012633A
Cited By
Hand-drawn style map generation system and method based on artificial intelligence and geographic information
CN121708155A
Image transmission and reception system, transmitter, receiver, computer program, and image transmission and reception method
JP7838169B1