Object position control method and device based on text graph model, equipment and medium
By binarizing and semantically encoding the initial image, extracting geometric features and performing text enhancement, the problem of inaccurate object positions in the text-based graph model is solved, and more efficient object position generation is achieved.
Patent Information
- Application Number
- CN202510763027.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2025-09-05
AI Technical Summary
Existing cultural image models are poor in controlling object positions and have a high degree of randomness, resulting in inaccurate object positions in the generated images, which cannot meet the needs of medical diagnosis and financial publicity.
By binarizing the initial image, the geometric features of the target object are extracted, and the target description text is generated by combining semantic encoding and text enhancement processing. The semantic feature vector is used to generate an accurate target object image.
The accuracy and efficiency of object position generation are improved, and images of target objects that are more in line with expectations can be generated, which reduces invalid calculations and improves image generation efficiency.
Smart Images

Figure CN120599042A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image detection technology, and in particular to an object position control method, device, equipment and medium based on a Wensheng graph model. Background Art
[0002] In text-to-image models, object position control relies on fuzzy positioning driven by semantic understanding, rather than precise geometric layout. Its core function is to provide "reasonable guesses that match the text description." Traditional text-to-image models can generate conceptual diagrams of object placement based on text descriptions (e.g., "In the kitchen, the cabinets are on the left, the stove is in the middle, and the refrigerator is on the right"), assisting in quickly visualizing layout solutions. However, current mainstream models, such as SDXL, have poor control over object position and exhibit significant randomness, significantly limiting output efficiency.
[0003] For example, in medical imaging diagnosis, doctors need to combine the text descriptions in the imaging report to accurately determine the location of the lesion. However, the traditional text-based image model has sparse and irregular descriptions of the object location. When describing a chest X-ray, the text may simply mention "there is a shadow in the lungs", but lack detailed descriptions of the specific location of the shadow in the lungs, such as the left upper lobe or the right middle lobe. As a result, the lesion location in the generated image is inaccurate and cannot meet the needs of medical diagnosis and teaching.
[0004] For example, in the promotion of financial products, it is necessary to generate promotional posters that contain specific financial elements (such as stock charts, currency symbols, financial product icons, etc.) and have a reasonable layout. Different financial practitioners or data labelers use non-standard terms and methods to describe the position of objects in the image. Most of them use relatively vague expressions such as "above" and "below", which leads to the traditional text-based graph model not being accurate enough in the object position when generating the corresponding image.
[0005] Therefore, how to improve the accuracy and efficiency of object position generation has become an urgent problem to be solved. Summary of the Invention
[0006] The present invention provides an object position control method, device, equipment and medium based on a cultural graph model, the main purpose of which is to solve the problems of inaccurate object position generation and low generation efficiency when generating an image of a target object.
[0007] In a first aspect, to achieve the above-mentioned objectives, the present invention provides an object position control method based on a Wensheng graph model, comprising:
[0008] Obtaining an initial image and initial position description text of the target object, and performing binarization processing on the initial image to obtain a target mask image;
[0009] Extracting geometric features of the target object in the target mask image, and determining object position data of the target object according to the geometric features;
[0010] Performing text enhancement processing on the initial position description text according to the object position data to obtain a target description text;
[0011] The target description text is semantically encoded to obtain a semantic feature vector, and a corresponding target object image is generated according to the semantic feature vector.
[0012] In a second aspect, the present invention further provides an object position control device based on a Wensheng graph model, comprising:
[0013] A binarization processing module is used to obtain an initial image and initial position description text of the target object, and perform binarization processing on the initial image to obtain a target mask image;
[0014] a position data calculation module, configured to extract geometric features of the target object in the target mask image and determine object position data of the target object based on the geometric features;
[0015] A position embedding text module is used to perform text enhancement processing on the initial position description text according to the object position data to obtain a target description text;
[0016] The object image generation module is used to perform semantic encoding on the target description text to obtain a semantic feature vector, and generate a corresponding target object image according to the semantic feature vector.
[0017] In a third aspect, the present invention further provides an electronic device, comprising:
[0018] at least one processor; and,
[0019] a memory communicatively connected to the at least one processor; wherein,
[0020] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can execute the object position control method based on the Wensheng graph model described above.
[0021] In a fourth aspect, the present invention further provides a computer-readable storage medium, wherein at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is executed by a processor in an electronic device to implement the above-mentioned object position control method based on the Wensheng graph model.
[0022] In an embodiment of the present invention, a thumbnail is obtained by scaling the initial image by a preset multiple, significantly reducing the amount of image data. This can reduce computing resource consumption and improve generation efficiency in subsequent processing. Grayscale processing simplifies image information while retaining key brightness features of the image, without affecting the distinction between the target object and the background, thereby improving the quality of image generation and the accuracy of subsequent position data acquisition. Geometric features such as shape and size can accurately describe the target object, providing a reliable basis for subsequent position calculation, facilitating the accurate control of the object position by the text-based graph model, and thus generating a more accurate target object image. Cleaning processing can remove redundant characters, incorrect expressions, and other noise in the initial position description text, improving text quality. Object position data is encoded into vectors, facilitating computer processing and storage of position information in a unified format, facilitating interaction with other data, and enriching the text content, making the position description more accurate and complete. Key semantic feature vectors are extracted through semantic coding to accurately capture the essential information of the target object, providing a reliable basis for position generation and ensuring that the object position in the generated image is more consistent with the expected one. When generating images using semantic feature vectors, combined with techniques such as spatial position constraints, the object position can be quickly located, reducing ineffective calculations and significantly improving image generation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0024] Figure 1 Schematic diagram of an application environment of an object position control method based on a Wensheng graph model in one embodiment of the present invention;
[0025] Figure 2 A schematic flow chart of an object position control method based on a Wensheng graph model provided in one embodiment of the present invention;
[0026] Figure 3 A schematic diagram of a process for semantically encoding the target description text provided in one embodiment of the present invention;
[0027] Figure 4 A schematic diagram of a module of an object position control device based on a Wensheng graph model provided by one embodiment of the present invention;
[0028] Figure 5 A schematic structural diagram of an electronic device for implementing an object position control method based on a Wensheng graph model provided by an embodiment of the present invention;
[0029] Figure 6Another structural diagram of an electronic device for implementing an object position control method based on a Wensheng graph model provided by an embodiment of the present invention.
[0030] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0031] In order to enable those skilled in the art to better understand the technical solutions of the present disclosure, and to fully understand and implement how the present disclosure applies technical means to solve technical problems and achieve the corresponding technical effects, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. The embodiments of the present disclosure and the various features in the embodiments can be combined with each other without conflict, and the technical solutions formed are all within the scope of protection of the present disclosure. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without making creative work should fall within the scope of protection of the present disclosure.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, apparatus, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0033] The embodiment of the present application provides an object position control method based on a Wensheng graph model, and the execution subject of the object position control method based on the Wensheng graph model includes but is not limited to a server, a terminal, and at least one of the electronic devices that can be configured to execute the device provided by the embodiment of the present application. In other words, the object position control method based on the Wensheng graph model can be executed by software or hardware installed on a terminal device or a server device. The server includes but is not limited to: a single server, a server cluster, a cloud server or a cloud server cluster, etc. The server can be an independent server, or it can be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0034] The object position control method based on the Wensheng graph model of the present invention can be applied to Figure 1 application environment. Among them, the client communicates with the server through the network. The server can obtain the initial image and initial position description text of the target object through the client, and obtain a thumbnail by scaling the initial image by a preset multiple, which greatly reduces the amount of image data. In subsequent processing, it can reduce computing resource consumption and improve generation efficiency. The grayscale processing simplifies the image information while retaining the key brightness features of the image, without affecting the distinction between the target object and the background, improving the quality of image generation and the accuracy of subsequent position data acquisition; geometric features such as shape and size can accurately describe the target object, provide a reliable basis for subsequent position calculation, and facilitate the cultural map model to accurately control the object position, thereby generating a more accurate target object image; cleaning processing can remove The redundant characters, incorrect expressions and other noises in the initial position description text are removed to improve the text quality. The object position data is encoded into a vector, which makes it easier for the computer to process and store the position information in a unified format, facilitates interaction with other data, and enriches the text content, making the position description more accurate and complete; through semantic coding, key semantic feature vectors are extracted to accurately capture the essential information of the target object, provide a reliable basis for position generation, and make the object position in the generated image more consistent with expectations; when using semantic feature vectors to generate images, combined with spatial position constraints and other technologies, the object position can be quickly located, invalid calculations can be reduced, and the image generation efficiency can be greatly improved. Finally, the target object image output is fed back to the client. Among them, the client can be but is not limited to various personal computers, laptops, smart phones, tablets and portable wearable devices. The server can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.
[0035] Reference Figure 2FIG. 1 is a flow chart of an object position control method based on a Wensheng graph model according to an embodiment of the present invention. In this embodiment, the object position control method based on a Wensheng graph model includes:
[0036] S1. Acquire an initial image and initial position description text of a target object, perform binarization processing on the initial image, and obtain a target mask image.
[0037] In an embodiment of the present invention, the initial image of the target object refers to the original image data of the target object in a specific scene, specifically refers to the original image that has not been extracted or processed specifically for the target object, and fully presents the visual information of the target object in the shooting environment, including the target object itself and the surrounding background, other objects, etc.
[0038] For example, in a street scene captured by a surveillance camera, if the target object is a red car, then the original image containing the car and the background such as the street, pedestrians, and buildings is the initial image of the target object.
[0039] In an embodiment of the present invention, the initial position description text is a textual description of the position of the target object in the image or actual scene, indicating the positional relationship of the target object relative to other objects, coordinate systems, or reference points. For example, for the red car mentioned above, the position description text may be "located on the right side of the street, near the entrance of a coffee shop."
[0040] In an embodiment of the present invention, the optical imaging system of the camera device can be used to focus the light reflected or emitted by the target object onto the image sensor, and the light signal is converted into a digital image signal through photoelectric conversion and signal processing, thereby obtaining an initial image of the target object.
[0041] Specifically, you can choose a suitable camera device, such as a digital camera, mobile phone camera, industrial camera, etc., determine the camera parameters such as focal length, aperture, shutter speed, sensitivity, etc. according to the characteristics of the target object and the shooting environment, adjust the camera position and angle, so that the target object is within the camera's field of view, and ensure that the image clearly and completely includes the target object. Press the shutter button, the camera completes the shooting and stores the image data in a storage medium, such as a memory card, mobile phone storage, etc., to obtain the initial image of the target object.
[0042] In the embodiment of the present invention, the binarization processing of the initial image to obtain the target mask image includes:
[0043] Scaling the initial image according to a preset magnification to obtain a thumbnail of the initial image;
[0044] Performing grayscale processing on the thumbnail to obtain a grayscale thumbnail;
[0045] Performing an inversion process on the grayscale thumbnail to obtain a first mask image;
[0046] Counting pixels of the first mask image to obtain the number of pixels;
[0047] performing pixel connectivity processing on the first mask image according to the number of pixels to obtain a connected domain area of the first mask image;
[0048] Identify the image region where the area of the connected domain is smaller than a preset binarization threshold as the background region of the initial image;
[0049] Identify the image region where the connected domain area is greater than the binarization threshold as the target object region of the initial image;
[0050] A target mask image of the initial image is generated according to the background area and the target object area.
[0051] In an embodiment of the present invention, image scaling is an operation of adjusting the image size by changing the size of the image. The preset multiple is usually a value less than 1, which is used to reduce the image, for example 0.5 times. The image processing library (such as the resize function in OpenCV) can be used to perform the scaling operation on the initial image. The algorithm will traverse each pixel of the initial image and calculate the position of the corresponding pixel in the thumbnail according to the scaling multiple to obtain the thumbnail.
[0052] Among them, the grayscale processing is the process of converting a color image into a grayscale image. Each pixel in the grayscale image contains only brightness information, not color information. For each initial pixel in the thumbnail, the values of its RGB (red, green, and blue) channels are obtained, and the grayscale value is calculated using the weighted average method. For example, it can be calculated using the formula Gray = 0.299*R+0.587*G+0.114*B. The calculated grayscale value is assigned to the initial pixel, thereby converting the color thumbnail into a grayscale thumbnail.
[0053] Among them, the inversion processing is an operation to invert the grayscale value of each pixel in the image, that is, black becomes white, white becomes black, and other grayscale values are also inverted accordingly. Each pixel in the grayscale thumbnail is traversed, the grayscale value of each pixel is inverted, and the inverted grayscale value is assigned to the grayscale pixel, thereby obtaining a first mask image, in which the originally darker area becomes brighter, and the brighter area becomes darker.
[0054] In detail, each pixel in the first mask image is traversed, the total number of all pixels in the image is counted, and pixel connectivity processing is performed on the first mask image according to the number of pixels. The first mask image can be processed using a connected region analysis algorithm, and a unique label is assigned to each connected region. The pixel position corresponding to each label is recorded, and the number of pixels in each connected region is counted to obtain the area of each connected domain.
[0055] Among them, the preset binary threshold is a threshold used to distinguish the target object and the background in the initial image. All connected domains are traversed, and the area of each connected domain is compared with the preset threshold. For a connected domain with an area smaller than the threshold, its corresponding area in the first mask image is marked as the background area. For a connected domain with an area larger than the threshold, its corresponding area in the first mask image is marked as the target object area. The target mask image is a binary image, in which the target object area is white (or a certain highlight color) and the background area is black (or a certain low-brightness color).
[0056] For example, in the field of medical health, the binarization processing method of the embodiment of the present invention can be applied to medical image analysis, such as melanoma detection in dermatoscope images; in dermatological diagnosis, doctors need to use a dermatoscope to take high-resolution images of the patient's skin surface to detect lesions such as melanoma. The initial image may contain complex structures such as hair, pigmentation, and blood vessels. The goal is to segment suspected lesion areas such as melanoma from the background.
[0057] In detail, the initial image is a color or grayscale image taken by a dermatoscope (usually in RGB format with a resolution of 1024×768 pixels or higher). The target object is the lesion area (such as melanoma), which may have irregular shapes, uneven colors, blurred edges and other features. The background is normal skin tissue, hair, blood vessels, pigmentation and other interferences.
[0058] Specifically, the original dermoscopic image is scaled by a preset multiple (such as 0.5 times) to generate a low-resolution thumbnail (such as 512×384 pixels) to reduce the amount of calculation. The pixel values of the grayscale thumbnail are inverted (such as changing 255 to 0 and 0 to 255) to reverse the contrast between the lesion area (usually darker) and the background (brighter) to facilitate subsequent segmentation. The value of each pixel in the inverted image is counted, and connected areas (such as pixel clusters in the lesion area) are identified; a preset binarization threshold (such as 500 pixels) is set, and areas with a connected domain area smaller than the threshold are marked as background (such as hair, blood vessels), and areas larger than the threshold are marked as target objects (such as melanoma). A binary mask image is generated based on the segmentation result, in which the target object area is white (255) and the background area is black (0).
[0059] Through automated segmentation, the embodiments of the present invention allow doctors to quickly locate suspected lesion areas, reducing manual labeling time. The mask image can be used for subsequent analysis (such as shape, color, and texture feature extraction), assisting in early screening of melanoma and improving the accuracy and efficiency of lesion detection.
[0060] In an embodiment of the present invention, a thumbnail is obtained by scaling the initial image by a preset multiple, which greatly reduces the amount of image data. In subsequent processing, the consumption of computing resources can be reduced, and processing efficiency can be improved. The grayscale processing simplifies the image information while retaining the key brightness features of the image, without affecting the distinction between the target object and the background, thereby improving the quality of image generation and the accuracy of subsequent position data acquisition.
[0061] S2. Extracting geometric features of the target object in the target mask image, and determining object position data of the target object according to the geometric features.
[0062] In an embodiment of the present invention, the geometric features are used to determine the specific position of the target object in the image or space. For example, the position coordinates of the minimum circumscribed rectangle of the target object clarify the starting position and ending position of the target object on the two-dimensional image plane. The coordinates of the geometric center point are also important geometric features. It represents the center position of the target object in a geometric sense and can be used to measure the relative position relationship of the target object in the overall scene.
[0063] In the embodiment of the present invention, extracting the geometric features of the target object in the target mask image includes:
[0064] Performing a connected domain analysis on the target mask image to determine at least one connected region containing the target object;
[0065] Generating a minimum rectangle for the connected area to obtain a minimum circumscribed rectangle, and calculating rectangle position coordinate parameters and size parameters of the minimum circumscribed rectangle;
[0066] Calculate the geometric center point coordinate parameters of the connected area according to the rectangle position coordinate parameters and size parameters;
[0067] The rectangle position coordinate parameters, the size parameters and the geometric center point coordinate parameters are used as geometric features of the target object.
[0068] In an embodiment of the present invention, connected domain analysis is a technology used to identify and mark interconnected pixel regions in an image. In a target mask image, by analyzing the connectivity between pixels, a set of pixels with the same attributes (such as the same grayscale value, usually white pixels representing the target object in the target mask image) and interconnected are divided into a connected region.
[0069] In detail, a marking matrix with the same size as the target mask image is initialized. Starting from the upper left corner of the image, a depth-first search (DFS) algorithm is used to traverse all adjacent white pixels in the target mask image, and the adjacent white pixels are marked with the same connected region number. The adjacent pixels of the adjacent pixels are then recursively or iteratively searched until no new white pixels can be added to the current connected region. The above process is repeated until all white pixels in the image are marked into the corresponding connected regions, thereby determining the connected region containing the target object.
[0070] Specifically, the minimum enclosing rectangle refers to a rectangle that can completely surround a given connected area and has the smallest area. All pixels in the connected area are traversed, and pixels at the edge of the connected area are recorded. The edge pixels constitute the boundary outline of the connected area. An algorithm such as the rotating calcaneal method can be used to generate the minimum enclosing rectangle. That is, for a given boundary outline, a rectangle is rotated so that it is always tangent to the boundary outline during the rotation process, and the rotation angle that can minimize the area of the rectangle and the corresponding rectangle vertex position are recorded.
[0071] Among them, the coordinates of the four vertices of the minimum circumscribed rectangle are determined to constitute the rectangle position coordinate parameters, and the size parameters include the width and height of the rectangle. The width can be obtained by calculating the difference between the maximum horizontal coordinate and the minimum horizontal coordinate of the rectangle in the horizontal axis direction, and the height can be obtained by calculating the difference between the maximum vertical coordinate and the minimum vertical coordinate in the vertical axis direction. In the case of a rectangle, the geometric center point is the intersection of the two diagonals of the rectangle, and the rectangle position coordinate parameters, the size parameters and the geometric center point coordinate parameters are used as the geometric feature representations of the target object.
[0072] In an embodiment of the present invention, determining the object position data of the target object according to the geometric features includes:
[0073] Calculating relative position data of the target object on the two-dimensional image plane according to the geometric features;
[0074] determining an initial bounding box position of the target object according to the geometric features, and performing center calibration on the initial bounding box position to obtain calibrated bounding box position data;
[0075] Performing coordinate transformation based on the calibration bounding box position data and preset spatial coordinate system parameters to obtain absolute position data of the target object in three-dimensional space;
[0076] The relative position data and the absolute position data are used as object position data of the target object.
[0077] In an embodiment of the present invention, key position information such as the coordinates of the geometric center point and the coordinates of the upper left corner of the minimum circumscribed rectangle are obtained from the extracted geometric features, and the geometric center point is used as a reference point for the relative position. The relative position data can then be expressed as the coordinate value of the geometric center point in the image coordinate system; for example, the horizontal and vertical coordinates of the geometric center point are recorded, and the horizontal and vertical coordinates represent the relative position of the target object in the horizontal and vertical directions of the image.
[0078] Specifically, the coordinates of the four vertices or the upper left corner and the width of the minimum bounding rectangle are obtained from the geometric features, the range of the initial bounding box is determined in the image, and it is checked whether the center of the initial bounding box coincides with the geometric center of the target object. If not, the position of the bounding box is adjusted so that its center is aligned with the geometric center of the target object; the adjustment method can be to translate the bounding box, and the translation distance is the coordinate difference between the center of the initial bounding box and the geometric center of the target object.
[0079] Among them, coordinate transformation uses the mapping relationship between the image coordinate system and the three-dimensional space coordinate system to transform the position information of the calibration bounding box on the two-dimensional image plane into the three-dimensional space. Specifically, a coordinate transformation model can be used to convert it into coordinates in the three-dimensional space, that is, absolute position data.
[0080] In an embodiment of the present invention, geometric features such as shape and size can accurately describe the target and provide a reliable basis for subsequent position calculation. The calculated position data includes two-dimensional relative and three-dimensional absolute positions, which facilitates the Wensheng graph model to accurately control the object position, thereby generating a more accurate target object image.
[0081] S3. Perform text enhancement processing on the initial position description text according to the object position data to obtain a target description text.
[0082] In an embodiment of the present invention, the object position data is converted into a position coding vector, and the initial position description text is denoised at the same time to obtain high-quality denoised text, and the denoised text is enhanced using the position coding vector to obtain the target description text.
[0083] In an embodiment of the present invention, performing text enhancement processing on the initial position description text according to the object position data to obtain the target description text includes:
[0084] Performing text cleaning on the initial position description text to obtain a cleaned text;
[0085] Encoding the object position data to obtain a position encoding vector;
[0086] The position encoding vector is embedded into the cleaned text to obtain a target description text.
[0087] In an embodiment of the present invention, the initial position description text may contain some irrelevant characters, special symbols, extra spaces or line breaks and other noise data, which affect the effect of subsequent text processing. These noises are identified and removed through text cleaning technology to make the text more standardized and pure.
[0088] In detail, traverse each character of the initial position description text and check whether the character belongs to a predefined set of irrelevant characters, such as punctuation marks (except the punctuation marks necessary in the position description), special symbols, etc. If the character belongs to the irrelevant character set, delete it from the text; at the same time, check whether there are multiple consecutive spaces or unnecessary line breaks in the text. For multiple consecutive spaces, replace them with one space; for unnecessary line breaks, delete or replace them according to the text format requirements.
[0089] Specifically, the object position data is discretized into a finite number of categories or intervals. For example, the coordinate range is divided into multiple intervals, each interval corresponds to a discrete value, and the discretized position data is converted into a vector form. Each element in the vector can represent a different feature or dimension of the position data. The vector representation can be easily integrated and processed with other text data.
[0090] Among them, according to business needs and data processing capabilities, the value range is divided into several intervals. For example, the horizontal axis range [0,100] is divided into 10 intervals, and the length of each interval is 10. For each position data, it is determined which interval it belongs to, and the interval number or identifier is used to represent a certain position data. According to the discretized position data, a fixed-length vector is constructed, and the interval number or identifier corresponding to the discretized position data is mapped to the corresponding position of the vector.
[0091] In detail, the position encoding vector is fused with the cleaned text so that the text contains position information. The position encoding vector can be regarded as the annotation information of certain positions in the text, and the vector information is associated with the text through the sequence labeling method; for example, the vector information can be embedded at the beginning, end or near the keywords related to the position description of the text.
[0092] For example, in the medical and health field, the text enhancement processing method of the embodiment of the present invention can be applied to medical imaging report generation or clinical document automation, such as the description of lesion location in chest X-ray (CXR).
[0093] For example, in radiology, doctors need to write reports for chest X-rays, describing the location and size of lesions (such as nodules and masses). This initial location description text may be generated by speech recognition or natural language processing (NLP) models, but lacks precise spatial positioning information. By incorporating the lesion's object location data (such as coordinates and area divisions), the initial text can be enhanced to generate a more accurate target description text.
[0094] Specifically, the initial location description text may be "The patient's chest X-ray shows a nodule in the right upper lobe of the lung, approximately 1.5 cm in size, with clear boundaries." This text does not specify the specific location of the nodule (e.g., which quadrant or anatomical region of the right upper lobe). A computer-aided detection (CADe) system automatically annotates the coordinates of the lesion (e.g., pixel coordinates or normalized coordinates), removes redundant information in the initial location description text (e.g., repeated descriptions of "nodule"), corrects grammatical errors (e.g., changing "approximately 1.5 cm in size" to "approximately 1.5 cm in diameter"), encodes the lesion coordinates into a vector, encodes the anatomical region into a one-hot vector, converts the location encoding vector and the anatomical region encoding into a natural language description, and inserts them into the cleaned text. Thus, the target description text may be "The lesion is located in the anterior segment of the right upper lobe of the lung, approximately 20% of the image height from the apex and approximately 30% of the image width from the midline."
[0095] The embodiment of the present invention embeds position coding, and the target description text clarifies the anatomical position and spatial relationship of the lesion, reduces ambiguity, and provides doctors with more detailed positioning information to assist in formulating surgical plans or follow-up strategies.
[0096] In an embodiment of the present invention, the cleaning process can remove redundant characters, incorrect expressions and other noise in the initial position description text, improve the text quality, make subsequent processing more accurate and efficient, encode the object position data into a vector, facilitate the computer to process and store position information in a unified format, facilitate interaction with other data, and enrich the text content, making the position description more accurate and complete.
[0097] S4. Perform semantic encoding on the target description text to obtain a semantic feature vector, and generate a corresponding target object image according to the semantic feature vector.
[0098] In an embodiment of the present invention, the target description text is segmented and mapped into an initial word vector sequence, and a dependency graph is generated through context association analysis. Based on this, the dependency vector is determined, the attention score is calculated, and then weighted encoding is performed to obtain a semantic feature vector. A preset image generator is used in combination with the semantic feature vector to generate an initial image, and a target image generator is obtained through loss calculation and parameter adjustment, and finally an image of the target object is generated.
[0099] like Figure 3As shown, in the embodiment of the present invention, the semantic encoding of the target description text to obtain a semantic feature vector includes:
[0100] Performing text segmentation on the target description text to obtain a segmentation sequence, and mapping each word in the segmentation sequence to a corresponding word vector to obtain an initial word vector sequence;
[0101] Performing context association analysis on the initial word vector sequence to construct a dependency graph;
[0102] Determining a dependency vector of each initial word vector in the initial word vector sequence according to the dependency graph;
[0103] Calculating an attention score between each of the initial word vectors and the dependency vector;
[0104] The initial word vector sequence is weighted feature encoded according to the attention score to obtain a semantic feature vector.
[0105] In an embodiment of the present invention, the forward maximum matching method starts from the beginning of the text and takes a character string as long as possible each time to match it with the words in the dictionary. If the match is successful, a word is split out. Otherwise, the character string length is reduced by one and the matching is continued until the match is successful or the character string length is 0.
[0106] The present invention can also input the target description text into the word segmentation tool based on word segmentation tools such as Jieba word segmentation, THULAC, etc. For example, the text "the object is on the left side of the table" and the word segmentation result is "object / on / table / left side", and an initial word vector sequence is obtained. For each word in the word segmentation sequence, the corresponding word vector is searched in the word vector model.
[0107] Specifically, models such as support vector machines (SVM) are used to train on a large-scale annotated dependency syntax tree library to learn the probability distribution of dependency relationships between words. When analyzing the initial word vector sequence, the model selects the most likely dependency center word and dependency relationship type for each word based on the learned probability distribution.
[0108] For example, the maximum entropy model calculates the probability of each dependency relationship based on the characteristics of the text (such as the part of speech of the vocabulary, distance, etc.), and then selects the dependency relationship with the highest probability as the analysis result to generate a dependency graph.
[0109] Specifically, the dependency graph is traversed to find the direct dependent words of each word (i.e., the dependency center word and the dependency subwords), and the word vectors of the direct dependent words are aggregated to obtain the dependency vector of the word. There are many ways to aggregate, such as simply averaging, weighted averaging, or splicing the word vectors of the dependent words. The attention score can be calculated by cosine similarity. Cosine similarity calculates the cosine value of the angle between two vectors. Its value range is between [-1,1]. The larger the value, the closer the directions of the two vectors are and the higher the similarity.
[0110] Among them, according to the calculated attention score, each word vector in the initial word vector sequence is weighted and summed. Specifically, each word vector is multiplied by the corresponding attention score, and then all weighted word vectors are added together to obtain the semantic feature vector; for example, if there are three word vectors, and their attention scores are 0.2, 0.5, and 0.3 respectively, then the semantic feature vector is the sum of these three word vectors multiplied by 0.2, 0.5, and 0.3 respectively. This method can suppress irrelevant information and thus better capture the semantics of the text.
[0111] In an embodiment of the present invention, generating a corresponding target object image according to the semantic feature vector includes:
[0112] Extracting spatial position constraint parameters of the semantic feature vector and performing masking on the spatial position constraint parameters to obtain a spatial attention mask;
[0113] Using a preset image generator to generate an image of the semantic feature vector according to the spatial attention mask to obtain an initial generated image;
[0114] Performing loss calculation on the initial generated image to obtain image loss, and adjusting parameters of the image generator according to the image loss to obtain a target image generator;
[0115] The target object image corresponding to the semantic feature vector is generated by using the target image generator.
[0116] In an embodiment of the present invention, a semantic feature vector is typically composed of features of multiple dimensions, each dimension may represent different semantic information. By analyzing the structure and meaning of the semantic feature vector, it is determined which dimensions are related to the spatial position of the object. Based on the analysis results, the dimensions related to the spatial position are selected from the semantic feature vector to form spatial position constraint parameters; for example, if there are multiple dimensions that represent the position offset of the object in the horizontal and vertical directions respectively, they can be combined into a two-dimensional position vector as the spatial position constraint parameter.
[0117] In detail, the spatial position constraint parameters can be masked according to a binary mask, where the binary mask has only two values, 0 and 1, and is used to simply select or exclude information of certain spatial positions; for example, a binary mask with the same size as the image is set, and the mask value corresponding to the position within a certain range of the image center is set to 1, and the other positions are set to 0.
[0118] Specifically, the image generator can be a generative adversarial network (GAN), which consists of a generator and a discriminator. The generator is responsible for generating an image based on the input semantic feature vector and spatial attention mask, and the discriminator is used to determine whether the generated image is real or generated. The semantic feature vector and the spatial attention mask are directly spliced together to form a longer vector, which is then used as the input of the image generator.
[0119] The present invention assigns different weights to different parts of the semantic feature vector based on the spatial attention mask, and then inputs the weighted semantic feature vector into the generator. The difference between the generated image and the real image at each pixel point is calculated, that is, the average of the square of the difference between the corresponding pixel values of the generated image and the real image is calculated. By calculating the image loss value and updating the parameters in the opposite direction of the image gradient, the convergence can be accelerated and the performance of the image generator can be improved.
[0120] Specifically, the target generator converts the semantic feature vector into image data based on the learned mapping relationship; for example, if the semantic feature vector describes a red circular object located in the center of the image, then the target image generator will generate an image that meets the description. During the generation process, the generator will comprehensively consider the semantic information and spatial position information to ensure that the generated image meets the semantic requirements and is reasonable in spatial position.
[0121] In the embodiments of the present invention, the accuracy and efficiency of object position generation can be significantly improved. Key semantic feature vectors are extracted through semantic coding, and the essential information of the target object is accurately captured, providing a reliable basis for position generation, so that the object position in the generated image is more in line with expectations. When using semantic feature vectors to generate images, combined with spatial position constraints and other technologies, the object position can be quickly located, invalid calculations can be reduced, and the image generation efficiency can be greatly improved.
[0122] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0123] like Figure 4 , which is a functional module diagram of an object position control device based on a Wensheng graph model provided by an embodiment of the present invention.
[0124] In the embodiment of the present disclosure, an object position control device based on a Wensheng graph model is provided. The object position control device based on a Wensheng graph model corresponds one-to-one with the object position control method based on a Wensheng graph model in the above embodiment. Figure 4 As shown, the object position control device 100 based on the Wensheng graph model can be installed in an electronic device. According to the functions to be implemented, the object position control device 100 based on the Wensheng graph model includes a binarization processing module 101, a position data calculation module 102, a position embedding text module 103, and an object image generation module 104. The functional modules are described in detail as follows:
[0125] A binarization processing module 101 is used to obtain an initial image and initial position description text of a target object, and perform binarization processing on the initial image to obtain a target mask image;
[0126] a position data calculation module 102, configured to extract geometric features of the target object in the target mask image and determine object position data of the target object based on the geometric features;
[0127] A position embedding text module 103 is configured to perform text enhancement processing on the initial position description text according to the object position data to obtain a target description text;
[0128] The object image generation module 104 is configured to perform semantic encoding on the target description text to obtain a semantic feature vector, and generate a corresponding target object image according to the semantic feature vector.
[0129] In one embodiment, when the binarization processing module 101 performs binarization processing on the initial image to obtain the target mask image, it is configured to:
[0130] Scaling the initial image according to a preset magnification to obtain a thumbnail of the initial image;
[0131] Performing grayscale processing on the thumbnail to obtain a grayscale thumbnail;
[0132] Performing an inversion process on the grayscale thumbnail to obtain a first mask image;
[0133] Counting pixels of the first mask image to obtain the number of pixels;
[0134] performing pixel connectivity processing on the first mask image according to the number of pixels to obtain a connected domain area of the first mask image;
[0135] Identify the image region where the area of the connected domain is smaller than a preset binarization threshold as the background region of the initial image;
[0136] Identify the image region where the connected domain area is greater than the binarization threshold as the target object region of the initial image;
[0137] A target mask image of the initial image is generated according to the background area and the target object area.
[0138] In one embodiment, when extracting the geometric features of the target object in the target mask image, the position data calculation module 102 is configured to:
[0139] Performing a connected domain analysis on the target mask image to determine at least one connected region containing the target object;
[0140] Generating a minimum rectangle for the connected area to obtain a minimum circumscribed rectangle, and calculating rectangle position coordinate parameters and size parameters of the minimum circumscribed rectangle;
[0141] Calculate the geometric center point coordinate parameters of the connected area according to the rectangle position coordinate parameters and size parameters;
[0142] The rectangle position coordinate parameters, the size parameters and the geometric center point coordinate parameters are used as geometric features of the target object.
[0143] In one embodiment, when determining the object position data of the target object according to the geometric features, the position data calculation module 102 is configured to:
[0144] Calculating relative position data of the target object on the two-dimensional image plane according to the geometric features;
[0145] determining an initial bounding box position of the target object according to the geometric features, and performing center calibration on the initial bounding box position to obtain calibrated bounding box position data;
[0146] Performing coordinate transformation based on the calibration bounding box position data and preset spatial coordinate system parameters to obtain absolute position data of the target object in three-dimensional space;
[0147] The relative position data and the absolute position data are used as object position data of the target object.
[0148] In one embodiment, when the position embedding text module 103 performs text enhancement processing on the initial position description text according to the object position data to obtain the target description text, it is configured to:
[0149] Performing text cleaning on the initial position description text to obtain a cleaned text;
[0150] Encoding the object position data to obtain a position encoding vector;
[0151] The position encoding vector is embedded in the cleaned text to obtain a target description text.
[0152] In one embodiment, when performing semantic encoding on the target description text to obtain a semantic feature vector, the object image generation module 104 is configured to:
[0153] Performing text segmentation on the target description text to obtain a segmentation sequence, and mapping each word in the segmentation sequence to a corresponding word vector to obtain an initial word vector sequence;
[0154] Performing context association analysis on the initial word vector sequence to construct a dependency graph;
[0155] Determining a dependency vector of each initial word vector in the initial word vector sequence according to the dependency graph;
[0156] Calculating an attention score between each of the initial word vectors and the dependency vector;
[0157] The initial word vector sequence is weighted feature encoded according to the attention score to obtain a semantic feature vector.
[0158] In one embodiment, when the object image generation module 104 generates the corresponding target object image according to the semantic feature vector, it is configured to:
[0159] Extracting spatial position constraint parameters of the semantic feature vector and performing masking on the spatial position constraint parameters to obtain a spatial attention mask;
[0160] Using a preset image generator to generate an image of the semantic feature vector according to the spatial attention mask to obtain an initial generated image;
[0161] Performing loss calculation on the initial generated image to obtain image loss, and adjusting parameters of the image generator according to the image loss to obtain a target image generator;
[0162] The target object image corresponding to the semantic feature vector is generated by using the target image generator.
[0163] In the present invention, the specific definitions of the object position control device based on the Vincent graph model can be found in the definitions of the object position control method based on the Vincent graph model above and will not be repeated here. Each module in the aforementioned object position control device based on the Vincent graph model can be implemented in whole or in part through software, hardware, or a combination thereof. Each of these modules can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each of these modules.
[0164] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of the object position control method based on the Wensheng graph model.
[0165] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of the object position control method based on the Wensheng graph model.
[0166] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0167] Obtaining an initial image and initial position description text of the target object, and performing binarization processing on the initial image to obtain a target mask image;
[0168] Extracting geometric features of the target object in the target mask image, and determining object position data of the target object according to the geometric features;
[0169] Performing text enhancement processing on the initial position description text according to the object position data to obtain a target description text;
[0170] The target description text is semantically encoded to obtain a semantic feature vector, and a corresponding target object image is generated according to the semantic feature vector.
[0171] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and apparatuses can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the module division is merely a logical function division, and actual implementation may employ other division methods.
[0172] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.
[0173] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.
[0174] In some implementations of this embodiment, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that when the computer program is executed by a processor, the steps of the method described in the above embodiment are implemented.
[0175] The readable storage medium of the present invention stores a computer program, which, when executed by a processor of an electronic device, can implement:
[0176] Obtaining an initial image and initial position description text of the target object, and performing binarization processing on the initial image to obtain a target mask image;
[0177] Extracting geometric features of the target object in the target mask image, and determining object position data of the target object according to the geometric features;
[0178] Performing text enhancement processing on the initial position description text according to the object position data to obtain a target description text;
[0179] The target description text is semantically encoded to obtain a semantic feature vector, and a corresponding target object image is generated according to the semantic feature vector.
[0180] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0181] The computer-readable storage medium may also store at least one computer-executable program / instruction, such as a computer-readable instruction. Computer-readable storage media include, but are not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Computer-readable storage media may include, for example, read-only memory (ROM), a hard disk, a flash memory, etc. For example, a non-transitory computer-readable storage medium may be connected to a computing device such as a computer, and then, when the computing device executes the computer-readable instructions stored on the computer-readable storage medium, the various methods described above may be performed.
[0182] In addition, the computer device may also include (but is not limited to) a data bus, an input / output (I / O) bus, a display, and input / output devices (eg, keyboard, mouse, speaker, etc.).
[0183] In one embodiment, the at least one computer executable instruction may also be compiled into or constitute a software product / computer program product, wherein one or more computer executable instructions are executed by a processor to perform the various functions and / or method steps in the embodiments described in the present technology.
[0184] Those skilled in the art will appreciate that all or part of the processes in the above-described embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-described methods. In particular, any reference to memory, storage, database, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory.
[0185] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0186] In the embodiments provided in the present disclosure, it should be understood that the disclosed devices and methods may also be implemented in other ways. The device embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a program segment or a part of a code, and the above-mentioned module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box may also occur in an order different from that marked in the accompanying drawings. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, may be implemented with a dedicated hardware-based system that performs the specified function or action, or may be implemented with a combination of dedicated hardware and computer instructions.
[0187] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.
[0188] It should be noted that if software tools or components other than those of our company appear in the embodiments of this application, they are only used for illustration and do not represent actual use.
Claims
1. The object position control method based on the Wensheng graph model is characterized by: The method comprises: Obtaining an initial image and initial position description text of the target object, and performing binarization processing on the initial image to obtain a target mask image; Extracting geometric features of the target object in the target mask image, and determining object position data of the target object according to the geometric features; Performing text enhancement processing on the initial position description text according to the object position data to obtain a target description text; The target description text is semantically encoded to obtain a semantic feature vector, and a corresponding target object image is generated according to the semantic feature vector.
2. The object position control method based on the Wensheng graph model according to claim 1, characterized in that: The binarization process is performed on the initial image to obtain a target mask image, comprising: Scaling the initial image according to a preset magnification to obtain a thumbnail of the initial image; Performing grayscale processing on the thumbnail to obtain a grayscale thumbnail; Performing an inversion process on the grayscale thumbnail to obtain a first mask image; Counting pixels of the first mask image to obtain the number of pixels; performing pixel connectivity processing on the first mask image according to the number of pixels to obtain a connected domain area of the first mask image; Identify the image region where the area of the connected domain is smaller than a preset binarization threshold as the background region of the initial image; Identify the image region where the connected domain area is greater than the binarization threshold as the target object region of the initial image; A target mask image of the initial image is generated according to the background area and the target object area.
3. The object position control method based on the Wensheng graph model according to claim 1, characterized in that: The extracting geometric features of the target object in the target mask image includes: Performing a connected domain analysis on the target mask image to determine at least one connected region containing the target object; Generating a minimum rectangle for the connected area to obtain a minimum circumscribed rectangle, and calculating rectangle position coordinate parameters and size parameters of the minimum circumscribed rectangle; Calculate the geometric center point coordinate parameters of the connected area according to the rectangle position coordinate parameters and size parameters; The rectangle position coordinate parameters, the size parameters and the geometric center point coordinate parameters are used as geometric features of the target object.
4. The object position control method based on the Wensheng graph model according to claim 1, characterized in that: The determining the object position data of the target object according to the geometric features includes: Calculating relative position data of the target object on the two-dimensional image plane according to the geometric features; determining an initial bounding box position of the target object according to the geometric features, and performing center calibration on the initial bounding box position to obtain calibrated bounding box position data; Performing coordinate transformation based on the calibration bounding box position data and preset spatial coordinate system parameters to obtain absolute position data of the target object in three-dimensional space; The relative position data and the absolute position data are used as object position data of the target object.
5. The object position control method based on the Wensheng graph model according to claim 1, characterized in that: The semantic encoding of the target description text to obtain a semantic feature vector includes: Performing text segmentation on the target description text to obtain a segmentation sequence, and mapping each word in the segmentation sequence to a corresponding word vector to obtain an initial word vector sequence; Performing context association analysis on the initial word vector sequence to construct a dependency graph; Determining a dependency vector of each initial word vector in the initial word vector sequence according to the dependency graph; Calculating an attention score between each of the initial word vectors and the dependency vector; The initial word vector sequence is weighted feature encoded according to the attention score to obtain a semantic feature vector.
6. The object position control method based on the Wensheng graph model according to claim 1, characterized in that: Generating a corresponding target object image according to the semantic feature vector includes: Extracting spatial position constraint parameters of the semantic feature vector and performing masking on the spatial position constraint parameters to obtain a spatial attention mask; Using a preset image generator to generate an image of the semantic feature vector according to the spatial attention mask to obtain an initial generated image; Performing loss calculation on the initial generated image to obtain image loss, and adjusting parameters of the image generator according to the image loss to obtain a target image generator; The target object image corresponding to the semantic feature vector is generated by using the target image generator.
7. The object position control method based on the Wensheng graph model according to claim 1, characterized in that: The performing text enhancement processing on the initial position description text according to the object position data to obtain the target description text includes: Performing text cleaning on the initial position description text to obtain a cleaned text; Encoding the object position data to obtain a position encoding vector; The position encoding vector is embedded into the cleaned text to obtain a target description text.
8. The object position control device based on the Wensheng graph model is characterized in that: The device comprises: A binarization processing module is used to obtain an initial image and initial position description text of the target object, and perform binarization processing on the initial image to obtain a target mask image; a position data calculation module, configured to extract geometric features of the target object in the target mask image and determine object position data of the target object based on the geometric features; A position embedding text module is used to perform text enhancement processing on the initial position description text according to the object position data to obtain a target description text; The object image generation module is used to perform semantic encoding on the target description text to obtain a semantic feature vector, and generate a corresponding target object image according to the semantic feature vector.
9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the object position control method based on the Wensheng graph model according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the object position control method based on the Wensheng graph model according to any one of claims 1 to 7 is implemented.