Image recognition method and apparatus, device, and storage medium

WO2026200270A1PCT designated stage Publication Date: 2026-10-01TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2026/075935
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2026-01-30
Publication Date
2026-10-01

Smart Images

  • Figure CN2026075935_01102026_PF_FP_ABST
    Figure CN2026075935_01102026_PF_FP_ABST
Patent Text Reader

Abstract

An image recognition method executed by a computer device. The method comprises: on the basis of entity categories associated with a reference image, performing entity detection in the reference image, and obtaining a plurality of candidate entity regions, each candidate entity region corresponding to one entity category and a confidence level of the entity category (401); screening out at least one first candidate region and at least one second candidate region from the plurality of candidate entity regions, a confidence level corresponding to the first candidate region being greater than a first threshold, and a confidence level corresponding to the second candidate region being not greater than the first threshold (402); segmenting, from each first candidate region, a first pixel region comprising an entity, and segmenting, from each second candidate region, a second pixel region comprising an entity (403); screening out at least one target pixel region from the obtained at least one second pixel region, each target pixel region accounting for a proportion of the second candidate region from which the target pixel region was segmented greater than a first proportion threshold (404); and, on the basis of the at least one first pixel region and the at least one target pixel region, obtaining a segmentation result of the reference image (405).
Need to check novelty before this filing date? Find Prior Art

Description

An image recognition method, apparatus, device, and storage medium

[0001] Related applications

[0002] This application claims priority to Chinese patent application filed on March 24, 2025, with application number 202510349435.5, entitled "An Image Recognition Method, Apparatus, Device and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0003] This application relates to the field of artificial intelligence technology, and in particular to an image recognition method, apparatus, device, and storage medium. Background Technology

[0004] In many application scenarios, there is a need to build 3D scenes with similar appearances and consistent styles based on 2D reference images. For example, in a user-generated content (UGC) game with a large asset library, based on a 2D reference image related to a forest uploaded by the user, the game searches the asset library for 3D models that match the entities (e.g., trees, rocks) in the 2D reference image; then, based on the positional relationship of each entity in the 2D reference image, the game drags the matching 3D models to specific locations to obtain the corresponding 3D scene game level.

[0005] To identify entities in a reference image, the relevant techniques first train an image segmentation model using a set of sample images. After training, the image segmentation model identifies the entity category corresponding to each pixel in the reference image and the confidence level of that entity category. Then, pixels with confidence levels less than a preset threshold are removed, and the pixels corresponding to the same entity category are merged to obtain the pixel regions of each entity in the reference image.

[0006] Since image segmentation models rely on a set of sample images for training, and the annotation of each sample image in the set typically only identifies the entity categories and corresponding locations of some important entities, when the trained image segmentation model is used to identify entities in a reference image, the confidence level of pixels corresponding to entity categories that were not labeled during training is often low. In this case, filtering pixels by removing those with confidence levels below a preset threshold will result in the removal of pixel regions of some entities in the reference image; in other words, this method will miss some pixel regions of entities in the reference image, thus reducing the accuracy of entity recognition. Summary of the Invention

[0007] This application provides an image recognition method, apparatus, device, and storage medium.

[0008] On one hand, embodiments of this application provide an image recognition method, which includes:

[0009] Based on the entity categories associated with the reference image, entity detection is performed in the reference image to obtain multiple entity candidate regions. Each entity candidate region corresponds to an entity category and the confidence level of that entity category.

[0010] At least one first candidate region and at least one second candidate region are selected from the plurality of entity candidate regions, wherein the confidence level of the first candidate region is greater than a first threshold and the confidence level of the second candidate region is not greater than the first threshold;

[0011] Each first candidate region is segmented to extract a first pixel region containing an entity; and each second candidate region is segmented to extract a second pixel region containing an entity.

[0012] At least one target pixel region is selected from at least one second pixel region obtained; wherein the proportion of each target pixel region in the corresponding second candidate region is greater than a first proportion threshold.

[0013] The segmentation result of the reference image is obtained based on the at least one first pixel region and the at least one target pixel region.

[0014] On one hand, embodiments of this application provide an image recognition device, which includes:

[0015] The detection module is used to perform entity detection in the reference image based on the entity categories associated with the reference image, and obtain multiple entity candidate regions, each entity candidate region corresponding to an entity category and the confidence level of that entity category;

[0016] The filtering module is used to filter at least one first candidate region and at least one second candidate region from the plurality of entity candidate regions, wherein the confidence level of the first candidate region is greater than a first threshold and the confidence level of the second candidate region is not greater than the first threshold.

[0017] The segmentation module is configured to segment a first pixel region containing an entity from each first candidate region; and to segment a second pixel region containing an entity from each second candidate region.

[0018] The filtering module is further configured to filter out at least one target pixel region from the obtained at least one second pixel region; wherein the proportion of each target pixel region in the corresponding second candidate region is greater than a first proportion threshold.

[0019] The recognition module is used to obtain the segmentation result of the reference image based on the at least one first pixel region and the at least one target pixel region.

[0020] On one hand, embodiments of this application provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described image recognition method.

[0021] On one hand, embodiments of this application provide a computer-readable storage medium storing a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform the steps of the image recognition method described above.

[0022] On one hand, embodiments of this application provide a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer device, cause the computer device to perform the steps of the above-described image recognition method.

[0023] Details of one or more embodiments of this application are set forth in the following drawings and description. Other features, objects, and advantages of this application will become apparent from the specification, drawings, and claims. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 is a schematic diagram of a system architecture provided in an embodiment of this application;

[0026] Figure 2 is a schematic diagram of an application scenario provided by an embodiment of this application;

[0027] Figure 3 is a schematic diagram of an application scenario provided by an embodiment of this application;

[0028] Figure 4 is a flowchart illustrating an image recognition method provided in an embodiment of this application;

[0029] Figure 5 is a flowchart illustrating a visual encoding process provided in an embodiment of this application;

[0030] Figure 6A is a flowchart illustrating a method for obtaining entity prompt words according to an embodiment of this application;

[0031] Figure 6B is a flowchart illustrating a method for obtaining entity prompt words according to an embodiment of this application;

[0032] Figure 6C is a flowchart illustrating a method for obtaining entity prompt words according to an embodiment of this application;

[0033] Figure 7 is a schematic diagram of an entity detection result provided in an embodiment of this application;

[0034] Figure 8A is a schematic diagram of an entity detection result provided in an embodiment of this application;

[0035] Figure 8B is a schematic diagram of an entity detection result provided in an embodiment of this application;

[0036] Figure 9A is a flowchart illustrating a quality verification and novelty verification process provided in an embodiment of this application;

[0037] Figure 9B is a flowchart illustrating an image recognition method provided in an embodiment of this application;

[0038] Figure 10 is a schematic diagram of the structure of an image recognition device provided in an embodiment of this application;

[0039] Figure 11 is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0040] To make the objectives, technical solutions, and beneficial effects of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0041] For ease of understanding, the terms used in the embodiments of this invention are explained below.

[0042] The embodiments of this application relate to artificial intelligence (AI) technology, and are mainly designed based on computer vision (CV) technology within artificial intelligence.

[0043] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0044] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0045] Computer vision is a science that studies how to enable machines to "see." More specifically, it refers to using cameras and computers to replace human eyes in recognizing and measuring targets, and then performing image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Large-scale modeling has brought significant changes to the development of computer vision technology; pre-trained models in various vision domains, after fine-tuning, can be quickly and widely applied to specific downstream tasks. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and other technologies, as well as common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0046] Transformer is a common deep learning model architecture widely used in natural language processing, computer vision (CV), and speech processing. Originally proposed as a sequence-to-sequence model architecture for machine translation, it consists of an encoder and a decoder. Both the encoder and decoder are composed of a series of identical Transformer blocks, each consisting of at least a multi-head self-attention layer and a feedforward neural network layer. Currently, Transformer has become a commonly used architecture in natural language processing and is frequently used as a pre-trained model. Besides language-related applications, Transformer is also applied in computer vision, audio processing, and other fields.

[0047] Embedding: It generally refers to the process of mapping data (such as text, images, audio, etc.) to a low-dimensional vector space. Specifically, embeddings can be seen as a feature representation learning technique that can transform raw data (such as words, sentences in text, or pixels in images) into low-dimensional, continuous vector representations that can capture the semantic or structural information of the original data.

[0048] Open Set Recognition (OSR) is an important research area in machine learning, aiming to solve the problem of identifying classes not seen during training (i.e., open set classes). Traditional classification models typically assume that test data only contains classes seen during training (closed set assumption), but in real-world applications, the model may encounter samples of unknown classes. The goal of open set detection models is not only to correctly classify samples of known classes but also to identify samples of unknown classes, adapting to the diversity of objects in the real world.

[0049] Full Segmentation Model: This is an image segmentation model in computer vision that aims to classify each pixel in an image, thereby achieving fine-grained image segmentation. Full segmentation models are commonly used for tasks such as semantic segmentation, instance segmentation, and panoptic segmentation.

[0050] The Segment Anything Model (SAM) is a comprehensive segmentation model designed to segment any object in an image without requiring task-specific training. It achieves zero-shot and few-shot segmentation capabilities by combining a powerful pre-trained model with flexible prompting mechanisms. In practical applications, image segmentation is performed using cues such as points, bounding boxes, or text.

[0051] An entity refers to any thing that has an independent identity or meaning, including physical objects, abstract concepts, events, etc. Physical objects can be visually recognizable objects such as tables, chairs, and cars, with a definite shape, size, and boundaries. In computer vision, objects are typically represented using bounding boxes or pixel-level masks.

[0052] Entity cues: These provide semantic information to the model through textual descriptions (such as category names, attribute descriptions, etc.), helping the model better understand and detect targets. Entity cues can provide semantic clues for unknown categories. For example, even if the model has not seen "kangaroo" in the training set, it may still detect the target through the description of the cue word "kangaroo." In the embodiments of this application, entity cues can be used as input to an open-set detection model to define the category of the entity to be detected. These entity cues are the foundation for the model to perform object detection and segmentation.

[0053] Object recognition, detection, and segmentation: core tasks of computer vision technology, aiming to understand the content of the entire image. Object recognition identifies the category of objects in an image, detection locates the position of objects and marks their bounding boxes, and segmentation extracts pixels belonging to the objects from the image.

[0054] Candidate region: refers to the initial detection bounding box obtained by the object detection task. It may contain redundancy and errors and needs further processing to determine the final detection bounding box. It generally contains the coordinates of several vertices.

[0055] Non-Maximum Suppression (NMS) is a post-processing technique in object detection tasks used to remove redundant bounding boxes and retain the most likely detection results. The core idea of ​​NMS is to retain only the bounding box with the highest confidence for each object, while suppressing other bounding boxes with high overlap with it.

[0056] User-generated content (UGC): In contrast to expert-generated content, it refers to the mode of content generation and creation by ordinary users, and is widely used in social media, games and online platforms.

[0057] Reference images for scene construction: Reference images used to build the scene, such as photos of bedrooms, conference rooms, etc., and construction-oriented images generated from the Internet and other systems. It is expected that the system will generate 2D or 3D scenes with similar appearance and function based on these reference images.

[0058] Main Entities, Decorative Entities, and Set Entities: Based on the importance of each entity in the scene construction, the entities in the reference image are divided into main entities and decorative entities. Main entities (such as tables, chairs, beds, etc.) generally have a larger area and a greater impact on scene construction; their absence or errors will significantly affect the effect. Decorative entities (such as small vases on a table) generally have a smaller area and a relatively smaller impact on scene construction; they can be filled using other methods. Set entities refer to entities that affect the overall layout (such as floors, doors, windows, lawns, water surfaces, etc.), providing the basic structure for the scene.

[0059] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, digital twins, virtual humans, robots, AI-generated content (AIGC), conversational interaction, smart healthcare, smart customer service, and game AI. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0060] The solutions provided in this application mainly involve the application of artificial intelligence technology in image recognition. In practical applications, by recognizing reference images used for scene building, the detection boxes, entity categories, and pixel regions of each entity in the reference image are obtained to support entity understanding and extraction from the reference image, providing a foundation for subsequent scene building tasks (such as retrieval, recommendation, and scene layout generation).

[0061] In the image recognition process, the scheme provided in this application takes a reference image and multiple entity categories associated with the reference image as input. Using an open set detection model, it performs target detection on the reference image based on multiple entity categories to obtain multiple entity candidate regions. From these candidate regions, a first candidate region with a confidence level greater than a first threshold is selected, while second candidate regions with a confidence level not greater than the first threshold are retained. An image segmentation model is used to segment a first pixel region containing an entity from each first candidate region; and a second pixel region containing an entity is segmented from each second candidate region. At least one target pixel region is selected from each of the obtained second pixel regions; wherein the proportion of the target pixel region in the corresponding second candidate region is greater than a first proportion threshold. Finally, by combining the obtained first pixel regions and target pixel regions, the segmentation result of the reference image is obtained.

[0062] It is understood that in the specific embodiments of this application, reference images and other related data are involved. When the embodiments in this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0063] The design concept of the embodiments of this application will be introduced below.

[0064] In many application scenarios, there is a need to build 3D scenes with similar appearances and consistent styles based on 2D reference images. For example, in a UGC game with a large library of assets, based on a 2D reference image related to a forest uploaded by the user, the system searches the asset library for 3D models that match the entities (e.g., trees, rocks) in the 2D reference image. Then, based on the layout of each entity in the 2D reference image, the system drags and scales the matching 3D models to specific positions to obtain the corresponding 3D scene game level. In the above scene building task, the system relies on the user to identify and understand the objects and layout in the reference image, and then search for relevant 3D models and drag and scale them to the appropriate positions. This results in a small indoor scene consisting of 10 objects taking about 1 hour to build, indicating low scene building efficiency.

[0065] Therefore, the goal is to improve the efficiency of scene construction tasks through AI technology. For example, AI can be used to identify, understand, detect, and segment objects, content, and layouts in reference images. Object detection and entity segmentation are crucial tasks in computer vision technology.

[0066] To identify entities in a reference image, the relevant techniques first train an image segmentation model using a set of sample images. After training, the image segmentation model identifies the entity category corresponding to each pixel in the reference image and the confidence level of that entity category. Then, pixels with confidence levels less than a preset threshold are removed, and the pixels corresponding to the same entity category are merged to obtain the pixel regions of each entity in the reference image.

[0067] Since image segmentation models rely on a set of sample images for training, and the annotation of each sample image in the set typically only identifies the entity categories and corresponding locations of some important entities, when the trained image segmentation model is used to identify entities in a reference image, the confidence level of pixels corresponding to entity categories that were not labeled during training is often low. In this case, filtering pixels by removing those with confidence levels below a preset threshold will result in the removal of pixel regions of some entities in the reference image; in other words, this method will miss some pixel regions of entities in the reference image, thus reducing the accuracy of entity recognition.

[0068] Secondly, in specific scene building tasks, entity recognition and segmentation need to adapt to open sets (i.e., potentially containing all object categories) and reference images with different domains (i.e., different styles, such as design concept art style and game / anime style). However, most traditional models are designed and trained on datasets of natural scenes, and only recognize entities of a few labeled categories, excluding non-salient entities in the entire image. This leads to a discrepancy between the actual image appearance style and the expected entity categories to be recognized in the reference image and the image appearance style and entity categories that the model can recognize, resulting in decreased performance in real-world scene building tasks.

[0069] In view of this, embodiments of this application provide an image recognition method, in which:

[0070] Using the entity categories associated with the reference image as clues, entity detection is performed in the reference image to obtain multiple entity candidate regions. The entity categories associated with the reference image are not limited to those labeled during the training phase, but can also include entity categories that were not labeled during the training phase. Therefore, when performing entity detection, entity candidate regions corresponding to entity categories labeled during the training phase, as well as entity candidate regions corresponding to entity categories that were not labeled, can be detected. This allows the obtained multiple entity candidate regions to cover as many of the entities to be identified as possible, thereby improving the comprehensiveness of entity detection.

[0071] Secondly, from multiple entity candidate regions, a first candidate region with a confidence level greater than a first threshold (i.e., a high-confidence candidate region) is selected, while a second candidate region with a confidence level not greater than the first threshold (i.e., a low-confidence candidate region) is retained. Then, each first candidate region is segmented into pixels to obtain a corresponding first pixel region, thus identifying the pixel regions of salient entities in the reference image. Simultaneously, each second candidate region is segmented into pixels to obtain a corresponding second pixel region, and a target pixel region with a larger proportion in each second candidate region is selected to identify the pixel regions of missed entities. Finally, based on the obtained first pixel regions and target pixel regions, the segmentation result of the reference image is generated. This not only identifies salient entities in the reference image but also uses low-confidence candidate regions to supplement undetected entities, greatly reducing the number of missed entities and thus improving the accuracy of entity recognition in the image.

[0072] In addition, in the scene building task, low-confidence entity candidate regions are retained and used to supplement entities that were not successfully detected. In this way, even if there are differences between the image style and labeled entity categories of the dataset used during training and the actual image style and expected entity categories of the reference images in the scene building task, entity omission problems can be detected and corrected in time, thereby improving the performance of the scene building task.

[0073] The following is a brief introduction to the system architecture diagram applicable to the technical solutions of the embodiments of this application. It should be noted that the system architecture diagram described below is only used to illustrate the embodiments of this application and is not intended to limit the scope of the application.

[0074] Referring to Figure 1, which is a system architecture diagram applicable to an embodiment of this application, the system architecture includes at least a terminal device 101 and a server 102. The number of terminal devices 101 can be one or more, and the number of servers 102 can also be one or more. This application does not specifically limit the number of terminal devices 101 and servers 102.

[0075] Terminal device 101 may be a smartphone, tablet computer, laptop computer, desktop computer, smart home appliance, smart voice interaction device, smart vehicle device, etc., but is not limited to these.

[0076] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, but it is not limited to these.

[0077] It should be noted that the method in this embodiment can be executed by the terminal device 101 or the server 102 alone, or by the terminal device 101 and the server 102 together.

[0078] When executed independently by terminal device 101, the open set detection model and the image segmentation model are deployed on terminal device 101. Terminal device 101 invokes the open set detection model to perform entity detection in the reference image based on the entity categories associated with the reference image, obtaining multiple entity candidate regions. Terminal device 101 selects at least one first candidate region and at least one second candidate region from the multiple entity candidate regions. The confidence level of the first candidate region is greater than a first threshold, and the confidence level of the second candidate region is not greater than the first threshold.

[0079] When executed independently by server 102, the open set detection model and the image segmentation model are deployed on server 102. Server 102 calls the open set detection model to perform entity detection in the reference image based on the entity categories associated with the reference image, obtaining multiple entity candidate regions. Server 102 selects at least one first candidate region and at least one second candidate region from the multiple entity candidate regions. The confidence level of the first candidate region is greater than a first threshold, and the confidence level of the second candidate region is not greater than the first threshold.

[0080] Terminal device 101 or server 102 invokes an image segmentation model to segment a first pixel region containing an entity from each first candidate region; and segments a second pixel region containing an entity from each second candidate region. At least one target pixel region is selected from each of the obtained second pixel regions; wherein the proportion of each target pixel region in the corresponding second candidate region is greater than a first proportion threshold; based on the obtained first pixel regions and target pixel regions, a segmentation result of the reference image is obtained.

[0081] When jointly executed by server 102 and terminal device 101, the open set detection model is deployed on terminal device 101, and the image segmentation model is deployed on server 102. Terminal device 101, based on the entity categories associated with the reference image, invokes the open set detection model to perform entity detection in the reference image, obtaining multiple entity candidate regions. It then selects at least one first candidate region and at least one second candidate region, where the confidence level of the first candidate region is greater than a first threshold, and the confidence level of the second candidate region is not greater than the first threshold. The selected first and second candidate regions are then sent to server 102.

[0082] Server 102 calls an image segmentation model to segment a first pixel region containing an entity from each first candidate region; and to segment a second pixel region containing an entity from each second candidate region. At least one target pixel region is selected from each of the obtained second pixel regions, wherein the proportion of each target pixel region in the corresponding second candidate region is greater than a first proportion threshold; and based on the obtained first pixel regions and target pixel regions, a segmentation result of the reference image is obtained.

[0083] Both server 102 and terminal device 101 may include one or more processors, memory, and interactive I / O interfaces. Furthermore, server 102 may be configured with a database to store model parameters of open set detection models and image segmentation models. The memory of server 102 and terminal device 101 may also store program instructions required for execution in the image recognition method provided in this application embodiment. These program instructions, when executed by the processor, can be used to implement the image recognition process provided in this application embodiment.

[0084] It should be noted that when the image recognition method provided in this application embodiment is executed by either server 102 or terminal device 101 alone, the system architecture of this application may also include only a single device, either server 102 or terminal device 101. Alternatively, server 102 and terminal device 101 may be considered as the same device. Of course, in practical applications, when the image recognition method provided in this application embodiment is executed by both server 102 and terminal device 101, server 102 and terminal device 101 may also be the same device. That is, server 102 and terminal device 101 may be different functional modules of the same device, or virtual devices virtualized by the same physical device.

[0085] In some embodiments, a user can initiate an image recognition process by providing a reference image and multiple entity prompts through a terminal device 101, where each entity prompt represents an entity category. The server 102 can then receive the reference image and multiple entity prompts provided by the user, determine the segmentation result of the reference image using the image recognition method of this embodiment, and return the result to the terminal device 101 for presentation.

[0086] In this embodiment, the terminal device 101 and the server 102 can communicate directly or indirectly through one or more networks. The network can be a wired network or a wireless network; for example, the wireless network can be a mobile cellular network or a Wireless-Fidelity (WIFI) network, or other possible networks. This embodiment does not limit the types of networks used.

[0087] The following is a brief introduction to the application scenarios to which the technical solutions of the embodiments of this application are applicable. It should be noted that the application scenarios described below are only for illustrating the embodiments of this application and are not intended to limit the scope. In specific implementation, the technical solutions provided by the embodiments of this application can be flexibly applied according to actual needs.

[0088] The solution provided in this application can be applied to scene building tasks in any scenario, such as game scenarios, virtual-real integration scenarios, shopping scenarios, and content recommendation scenarios. This solution can be used in conjunction with various upstream and downstream systems based on scene building. Furthermore, this solution can be used as a foundational technology in various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, and audio-visual scenarios.

[0089] The following provides an illustrative example of the application scenarios to which the solution provided in this application is applicable.

[0090] Application Scenario 1: Generating UGC game levels in a game setting.

[0091] The system uploads a reference image based on the object being operated on, or generates a reference image in conjunction with an upstream reference image generation system. Then, it uses the scheme in this application embodiment to identify the main objects and layouts in the reference images. Next, it reconstructs these main objects in the game environment and scales and moves them according to the identified layouts to generate corresponding 2D or 3D UGC game level images.

[0092] For example, referring to Figure 2, the user uploads a photo 201 of a forest. The image recognition scheme in this application is used to recognize the photo 201, and a recognition result 202 is obtained. The recognition result 202 includes: detection boxes, pixel regions, and entity categories of objects such as trees and rocks in the photo 201.

[0093] In the downstream system, the image search of the material library is performed according to the recognition result 202 to obtain entity materials that match objects such as trees and rocks. Based on the position information of each object in photo 201, the corresponding entity materials are dragged and scaled to specific positions in the game environment, and then layout and scene optimization are performed to obtain the corresponding 2D scene game level image 203.

[0094] Application Scenario 2: Virtual-Real Integration Scenario.

[0095] Referring to Figure 3, to better demonstrate smart home technology, a home decoration design image 301 is first uploaded. The image recognition scheme in this application is used to recognize the home decoration design image 301, and a recognition result 302 is obtained. The recognition result 302 includes: the detection box, pixel area, and entity category of each piece of furniture.

[0096] In conjunction with the downstream system, an image search is performed on the 3D material library based on the recognition result 302 to obtain a matching 3D model for each piece of furniture. 3D estimation is performed on the home decoration design image 301 to obtain the depth information of each piece of furniture. Based on the positional relationship and depth information of each piece of furniture in the home decoration design image 301, the corresponding 3D models of each piece of furniture are placed, and layout and scene optimization are performed to obtain an interactive 3D scene image 303 for home decoration design, thereby helping to better showcase the 3D effects of each piece of furniture.

[0097] It should be noted that the embodiments of this application are not limited to the application scenarios exemplified above, but may also be other application scenarios, which are not specifically limited in this application.

[0098] Based on the system architecture diagram shown in Figure 1, this application embodiment provides a flowchart of an image recognition method, as shown in Figure 4. The flowchart of this method is executed by a computer device, which may be the terminal device 101 and / or the server 102 shown in Figure 1, and includes the following steps:

[0099] Step 401: Based on the entity categories associated with the reference image, perform entity detection in the reference image to obtain multiple entity candidate regions.

[0100] Specifically, the reference images in this application can be reference images used in scene building tasks, and these reference images can come from the Internet, be generated by other systems, or be input from objects being manipulated. Examples include downloaded photos of meeting rooms or taken photos of bedrooms. Of course, reference images can also be images used in other application scenarios; this application does not specifically limit their use.

[0101] Each entity category associated with the reference image can be represented by category hint text, which includes multiple entity hint words, each entity hint word corresponding to an entity category. Of course, in addition to multiple entity hint words, the category hint text may also include other descriptive text associated with the reference image, which is not specifically limited in this application.

[0102] After obtaining the category hint text, the category hint text and reference image are input into the open set detection model for entity detection (i.e., object detection), resulting in multiple entity candidate regions. Entity candidate regions are preliminary detection boxes obtained from the entity detection task, and may contain redundancy and errors. Each entity candidate region corresponds to an entity category and a confidence level for that entity category. The confidence level represents the probability that an entity within the entity candidate region belongs to the entity category corresponding to that entity candidate region; the higher the confidence level, the more accurate the entity detection result.

[0103] In some embodiments, the present application uses at least the following methods to obtain entity hint words corresponding to each entity category:

[0104] Feature extraction is performed on the reference image to obtain corresponding image features. Then, based on the image features, the scene category and multiple preliminary descriptive words of the reference image are obtained. From multiple preset word libraries, a target word library matching the scene category is selected. The target word library includes multiple supplementary descriptive words. Each supplementary descriptive word describes an entity whose frequency of occurrence is greater than a preset threshold under that scene category. Then, based on the multiple preliminary descriptive words and multiple supplementary descriptive words, multiple entity prompt words are obtained. The preset threshold can be determined according to different scene categories and actual data statistics. For example, in general scenes, the preset threshold can be set to 0.1-0.3 (i.e., frequency of occurrence greater than 10%-30%).

[0105] In this embodiment, the reference image is visually encoded using a visual encoder in an image description model to obtain image features. The image description model can be a large language model or other deep learning recognition model. In practical applications, the visual encoder can be implemented using a Vision Transformer (VIT), or it can be the encoding part of other possible models; this embodiment does not limit this approach.

[0106] Referring to Figure 5, this is a flowchart illustrating a visual encoding process provided in an embodiment of this application. First, the reference image is segmented to obtain multiple image blocks. Specifically, a uniform grid segmentation method can be used to divide the reference image into image blocks of the same size. Assuming the size of the reference image is H×W (height H, width W), and the size of each image block is h×w, then the number of rows in the image block is... The number of columns is in This indicates rounding down to the nearest integer.

[0107] Image patches are input into a vectorization layer, which performs image encoding processing on multiple image patches separately to obtain corresponding image patch encoding features. Specifically, in the vectorization layer, an embedding operation is performed on each image patch to convert it into a low-dimensional and continuous vector representation. This vector representation is the image patch encoding feature, which captures the semantic or structural information of the original image patch. Assuming each image patch is a three-dimensional tensor of length h×w×c (where h and w are the height and width of the image patch, respectively, and c is the number of channels), it is flattened into a one-dimensional vector of length h×w×c. Then, a fully connected layer maps this one-dimensional vector to a low-dimensional space to obtain a low-dimensional vector representation. Let the weight matrix of the fully connected layer be... The bias vector is Where d is the dimension of the low-dimensional space, the image block encoding feature e can be expressed as e = W·flatten(image_block) + b, where flatten(image_block) represents the flattening operation on the image block.

[0108] Based on the positions of multiple image patches in the image data, corresponding positional coding features are obtained. Then, based on these positional coding features, the resulting image patch coding features are serialized to obtain the image features of the reference image. Specifically, for each image patch, positional coding features are added to its corresponding image patch coding features based on its position in the image data. Then, the positional coding features and the image patch coding features corresponding to each image patch are concatenated to obtain the concatenated coding features. The concatenated coding features of multiple image patches constitute a coding feature sequence. This coding feature sequence is then fed into the coding layer for serialization coding to obtain the corresponding image features.

[0109] The encoding layer can be the encoder in a Transformer model, which comprises multiple identical extraction layers. Each extraction layer includes: a multi-head self-attention mechanism layer, a feedforward neural network, residual connections, and a layer normalization layer. The multi-head self-attention mechanism layer is used to capture global dependencies in the sequence; the feedforward neural network is used for non-linear transformations to enhance the model's expressive power; the residual connections and layer normalization layers are used to: directly add the input to the output of the sub-layer to alleviate the vanishing gradient problem, help the model train deeper networks, and accelerate training and improve model stability through normalization operations.

[0110] After the encoded feature sequence is input into the encoder, each extraction layer in the encoder performs the following operations in sequence: 1. Multi-head self-attention mechanism operation; 2. Residual connection and layer normalization; 3. Feedforward neural network transformation operation; 4. Residual connection and layer normalization. This encoder transforms the encoded feature sequence into image features of a reference image in a high-dimensional representation through multiple extraction layers.

[0111] Specifically, for each extraction layer, a multi-head self-attention mechanism is first performed, assuming the input encoded feature sequence is... (N is the sequence length, d is the feature dimension), the output Z1 of the multi-head self-attention mechanism can be calculated through the following steps: First, X is compared with the three weight matrices respectively. and Multiplying them yields the query matrix Q, the key matrix K, and the value matrix V, i.e., Q = XW. Q K = XW K V = XW V Then calculate the attention score. Finally, the output Z1 = SV of the multi-head self-attention mechanism is obtained. Next, residual connections and layer normalization are performed, resulting in the output Z2 = LayerNorm(X + Z1). Then, a feedforward neural network transformation is performed. The feedforward neural network consists of two fully connected layers. Let the weight matrix of the first fully connected layer be... The bias vector is The weight matrix of the second fully connected layer is The bias vector is The output Z3 of the feedforward neural network is Z3 = W2·ReLU(W1Z2+b1)+b2. Finally, a residual connection and layer normalization are performed, resulting in the output Z4 = LayerNorm(Z2+Z3). After processing through multiple extraction layers, the final image features are obtained.

[0112] It should be noted that the encoding layer in this application embodiment may also adopt other possible encoder structures, and this application embodiment does not limit this.

[0113] In one embodiment, the corresponding positional coding features can be obtained by using sinusoidal positional coding based on the positions of multiple image patches in the image data. Specifically, for position (i,j) in a two-dimensional image, its positional coding feature PE (i,j),2k and PE (i,j),2k+1 It can be calculated using the following formula: (For the i-direction); (For the j-direction), where d is the dimension of the position encoding, and k is the encoding index, with a value range of...

[0114] Next, based on the image features of the reference image, the image description model obtains the scene category and several preliminary descriptive terms of the reference image. The scene category of the reference image represents the scene corresponding to the reference image as a whole; for example, indoor or outdoor. Each preliminary descriptive term is used to describe an entity initially identified from the reference image; for example, house, flower shed, signboard, water, etc.

[0115] In practical applications, the preset thesaurus is also called the hot word library or high-frequency word library. Among multiple preset thesauruses, each preset thesaurus corresponds to a scene category, and each preset thesaurus includes multiple supplementary descriptive words. Each supplementary descriptive word is used to describe an entity that appears more frequently than a preset threshold in the corresponding scene category; each supplementary descriptive word is a common and easily overlooked descriptive word in the corresponding scene category.

[0116] The system selects a target vocabulary that matches the scene category of the reference image from multiple preset vocabulary lists, and then obtains multiple supplementary descriptive words from the target vocabulary list. These initial and supplementary descriptive words can be used directly as entity prompts; alternatively, preprocessing can be performed on the initial and supplementary descriptive words to obtain multiple entity prompts.

[0117] For example, referring to Figure 6A, the reference image 601 is input into the image description model to obtain the scene category "outdoor" and several preliminary descriptive terms "house, flower pot, signboard, water".

[0118] Obtain a target word library that matches the scene category "outdoor" from multiple hot word libraries, and obtain multiple supplementary descriptive words from the target word library, namely: lawn, street lamp, fence, potted plant, flower pot.

[0119] If multiple initial descriptors and multiple supplementary descriptors are used as entity prompts, the resulting entity prompts include: house, flower pot, signboard, water, lawn, street lamp, fence, potted plant, flower pot.

[0120] In this embodiment, an image description model is used to automatically generate preliminary descriptive words corresponding to the reference image as entity prompts, improving the efficiency of prompt acquisition. Secondly, since different hot word libraries contain common prompts and easily overlooked prompts in different scenarios, matching the reference image to the corresponding hot word library and supplementing the entity prompts of the reference image with prompts from that library ensures that necessary prompts are added in different scenarios, making the entity prompts of the reference image more comprehensive.

[0121] In some embodiments, in addition to obtaining entity prompts through image description models and hot word libraries, entity prompts are also obtained based on external input, specifically including the following operations:

[0122] Obtain the topic descriptor words set for the corresponding reference image; then, based on the semantic information contained in the topic descriptor words, perform text expansion operations on the topic descriptor words to obtain multiple target descriptor words; then, based on the multiple target descriptor words, multiple preliminary descriptor words, and multiple supplementary descriptor words, obtain multiple entity prompt words.

[0123] Specifically, a topic descriptor for the reference image is set based on external input. The topic descriptor is used to describe the topic of the reference image. The external input can be upstream input, such as input from the object being manipulated or other system inputs.

[0124] The externally input topic description words are used to expand the text of the prompt word expansion model, resulting in multiple target prompt words. The prompt word expansion model can be a large language model or other deep learning models.

[0125] Specifically, in the prompt word expansion model, features are first extracted from the topic description words to obtain corresponding text features, thereby capturing the semantic information contained in the topic prompt words. Word embedding can be used to extract features from the topic description words. Assuming the topic description words consist of m words, a pre-trained word embedding model (such as Word2Vec, GloVe, etc.) is used to map each word to a d-dimensional vector. Let the vector representation of the i-th word be v. i Then the text features of the topic descriptors can be the average of these word vectors, i.e. Alternatively, features can be obtained by concatenating and further processing them through a fully connected layer. Then, based on the semantic information contained in the topic descriptors, text expansion is performed on the topic descriptors to obtain multiple richer and more detailed target descriptors. The cue word expansion model can be a fine-tuning model based on a pre-trained language model (such as the GPT series). The topic descriptors are input into the model, and the model generates a series of texts semantically related to the topic descriptors based on the pre-trained language knowledge and the information learned in the fine-tuning stage. For example, the model will generate related modifiers, synonyms, hyponyms, etc., based on the keywords in the topic descriptors. By setting the maximum length of the generated text and some constraints (such as avoiding the generation of duplicate words), multiple target descriptors can be obtained.

[0126] Multiple target descriptors, multiple preliminary descriptors, and multiple supplementary descriptors can be directly used as entity prompts; alternatively, preprocessing operations can be performed on the multiple target descriptors, multiple preliminary descriptors, and multiple supplementary descriptors to obtain multiple entity prompts.

[0127] For example, referring to Figure 6B, the reference image 601 is input into the image description model to obtain the scene category "outdoor" and several preliminary descriptive terms "house, flower pot, signboard, water".

[0128] Obtain a target word library that matches the scene category "outdoor" from multiple hot word libraries, and obtain multiple supplementary descriptive words from the target word library, namely: lawn, street lamp, fence, potted plant, flower pot.

[0129] The system receives the topic descriptor "small wooden house by the sea" from the upstream input reference image. It then inputs this topic descriptor "small wooden house by the sea" into the prompt word expansion model to obtain multiple target descriptors for the reference image, namely: beautiful, wooden house, beach, and tree.

[0130] If the multiple target descriptors, multiple preliminary descriptors, and multiple supplementary descriptors obtained are all used as entity prompts, then the multiple entity prompts obtained include: house, flower pot, signboard, water, lawn, street lamp, fence, potted plant, flower pot, beautiful, wooden house, beach, tree.

[0131] In this embodiment, the subject descriptive words of the reference image are set through external input, and these subject descriptive words are expanded to obtain more descriptive words. When the entity cues of the reference image are supplemented with the expanded descriptive words, the flexibility of the cue source is improved, making the entity cues of the reference image more comprehensive.

[0132] In some embodiments, to improve the accuracy of the obtained entity prompts, this application performs preprocessing operations on the obtained multiple target descriptors, multiple preliminary descriptors, and multiple supplementary descriptors to obtain multiple entity prompts, specifically including the following operations:

[0133] Based on multiple target descriptors, multiple preliminary descriptors, and multiple supplementary descriptors, a descriptor set is constructed; then, a preprocessing operation is performed on the descriptor set, and multiple descriptors in the preprocessed descriptor set are used as multiple entity prompt words;

[0134] The preprocessing operations include at least one of the following: deduplication of prompt words; removal of erroneous words; supplementation of synonyms; replacement of ambiguous words; and supplementation of invalid words.

[0135] Specifically, the deduplication operation refers to removing duplicate prompts from the descriptive word set to avoid duplicate detection and recognition.

[0136] Error word removal refers to removing error words or abstract words that are irrelevant to the image recognition task in order to avoid generating overly broad detection boxes or unclear objects, such as abstract words like "beautiful" or "high quality".

[0137] In practical applications, an error word list can be pre-set, containing frequently occurring error words. Then, for each descriptor to be processed in the descriptor set, the descriptor is compared with each error word in the error word list. If the error word list contains an error word that matches the descriptor, the descriptor is removed from the descriptor set; if the error word list does not contain an error word that matches the descriptor, the descriptor is retained.

[0138] Of course, different error word lists can be set for different scene categories, with each error word list containing frequently occurring error words in the corresponding scene category. After obtaining the scene category of the reference image, a matching error word list is obtained from multiple error word lists based on the scene category of the reference image. Then, the matching error word list is used to remove error words from the descriptive word set. The specific removal process has been described above and will not be repeated here.

[0139] Synonym supplementation refers to adding synonyms or related words to all or part of the descriptive words in the descriptive word set to improve the accuracy of the descriptive words and avoid missed detections. For example, if the descriptive word "flowerpot" exists in the descriptive word set, synonyms such as "potted plant" and "vase" can be added to achieve secondary entity differentiation during image recognition.

[0140] In practical applications, a synonym list can be pre-set, containing multiple synonym groups, each of which includes several descriptive words with similar or related meanings. Then, for each descriptive word to be processed in the descriptive word set, the descriptive word is compared with each synonym group in the synonym list. When the synonym list contains a matching synonym group for the descriptive word, all other descriptive words in that synonym group are added to the descriptive word set.

[0141] Of course, different synonym lists can be set up for different scene categories, with each list containing frequently occurring synonym phrases for the corresponding scene category. After obtaining the scene category of the reference image, a matching synonym list is obtained from multiple synonym lists based on the scene category of the reference image. Then, the matching synonym list is used to supplement all or part of the synonyms of the descriptive words in the descriptive word set. The specific supplementation process has been introduced earlier and will not be repeated here.

[0142] Additionally, a commonly used thesaurus can be pre-set. For each descriptor to be processed in the descriptor set, the text similarity between that descriptor and each candidate word in the thesaurus is calculated. Candidate words with text similarity greater than a preset threshold are added to the descriptor set as synonyms. The text similarity can be calculated using Levenshtein Distance, which refers to the minimum number of editing operations (insertion, deletion, replacement) required to transform one string into another. Let the descriptor be s1, the candidate word be s2, and the edit distance be d, then the text similarity... Where |s1| and |s2| are the lengths of s1 and s2, respectively. Candidate words with text similarity greater than a preset threshold are added to the descriptor set as synonyms of the descriptor.

[0143] Ambiguity word replacement refers to replacing or supplementing descriptive words in a descriptive word set that are prone to ambiguity, in order to reduce post-processing errors caused by ambiguous words. For example, replacing the descriptive word "water" with "bottle" or "lake" can adapt to different scenarios, thereby ensuring more accurate entity recognition.

[0144] In practical applications, a replacement word list can be pre-set, containing multiple replacement word groups. Each replacement word group includes an ambiguous word and at least one corresponding replacement word. Then, for each descriptor to be processed in the descriptor set, the descriptor is compared with each ambiguous word in the replacement word list. When the replacement word list contains an ambiguous word that matches the descriptor, at least one corresponding replacement word is used to replace the descriptor in the descriptor set.

[0145] Of course, different replacement word lists can be set for different scene categories. Each replacement word list contains frequently occurring ambiguous words in the corresponding scene category and their corresponding replacement words. After obtaining the scene category of the reference image, a matching replacement word list is obtained from multiple replacement word lists based on the scene category of the reference image. Then, the matching replacement word list is used to replace or supplement the ambiguous descriptive words in the descriptive word set. The specific supplementation process has been introduced earlier and will not be repeated here.

[0146] Invalid word supplementation refers to the process where, during entity recognition, some entities in the reference image do not need to be identified. These entities are also called invalid entities irrelevant to the recognition task. For example, shadows in the reference image. The entity hints corresponding to invalid entities are called invalid words. In the actual recognition process, if only the entity hints (i.e., valid words) of valid entities are input without inputting invalid words, invalid entities will be output according to the entity category indicated by the valid words, thus introducing bias and errors.

[0147] For example, if the invalid word "shadow" is not input, the "shadow" in the reference image will be output as the valid word "door," resulting in a false positive. Therefore, it is necessary to add common invalid words to the descriptor set in advance.

[0148] You can pre-set an invalid word list containing common invalid words, and then add all or part of the invalid words in the invalid word list to the descriptor set.

[0149] Of course, different invalid word lists can be set for different scene categories, with each list containing frequently occurring invalid words in the corresponding scene category. After obtaining the scene category of the reference image, a matching invalid word list is retrieved from multiple invalid word lists based on the scene category of the reference image. All or part of the invalid words in the matching invalid word list are added to the descriptive word set.

[0150] In some embodiments, when preprocessing the descriptor set using multiple preprocessing operations such as deduplication of prompt words, removal of erroneous words, synonym supplementation, ambiguous word replacement, and invalid word supplementation, the execution order of these preprocessing operations can be preset. Different execution orders may affect the final entity prompt words. For example, performing error word removal first can avoid introducing unnecessary erroneous words in subsequent operations such as synonym supplementation. Generally, it is recommended to perform deduplication of prompt words and removal of erroneous words first to reduce the amount of data and error interference in subsequent processing, followed by synonym supplementation, ambiguous word replacement, and invalid word supplementation. Multiple preprocessing operations are executed sequentially according to the preset execution order to obtain the final descriptor set, and all descriptors in the final descriptor set are used as entity prompt words for subsequent entity detection. For example, the execution order of the multiple preprocessing operations can be set as: deduplication of prompt words, removal of erroneous words, synonym supplementation, ambiguous word replacement, and invalid word supplementation.

[0151] It should be noted that the execution order of multiple preprocessing operations is not limited to the one mentioned above, and can also follow other preset execution orders; in the actual preprocessing process, all or part of the above multiple preprocessing operations can be executed repeatedly, and this application does not make specific limitations on this.

[0152] Furthermore, when performing preprocessing operations on multiple initial descriptors and multiple supplementary descriptors to obtain multiple entity prompts, a descriptor set is also constructed based on the multiple initial descriptors and multiple supplementary descriptors; then, preprocessing operations are performed on the descriptor set, and multiple descriptors in the preprocessed descriptor set are used as multiple entity prompts. The preprocessing operations have been described earlier and will not be repeated here.

[0153] For example, referring to Figure 6C, the reference image 601 is input into the image description model to obtain the scene category "outdoor" and several preliminary descriptive terms "house, flower pot, signboard, water".

[0154] Obtain a target word library that matches the scene category "outdoor" from multiple hot word libraries, and obtain multiple supplementary descriptive words from the target word library, namely: lawn, street lamp, fence, potted plant, flower pot.

[0155] The topic descriptor "small wooden house by the sea" of the reference image input by the operation object is obtained. The topic descriptor "small wooden house by the sea" is input into the prompt word expansion model to obtain multiple target descriptors of the reference image, namely: beautiful, wooden house, beach, tree.

[0156] The aforementioned preliminary descriptive words, supplementary descriptive words, and target descriptive words constitute the initial descriptive word set, which includes: house, flower pot, signboard, water, lawn, street lamp, fence, potted plant, flower pot, beautiful, wooden house, beach, and tree.

[0157] First, perform a deduplication operation on the descriptor set, that is, remove duplicate descriptors from the descriptor set; for example, if the descriptor set contains two identical descriptors "flower pot", then remove one of the descriptors "flower pot" from the descriptor set.

[0158] Next, an error word removal operation is performed, which involves reading an error word list from the database. This list includes error words such as "beautiful," "scene," and "high quality." Then, this error word list is used to remove error words from the descriptive word set. For example, the descriptive word "beautiful" is removed from the descriptive word set.

[0159] Then, a synonym supplementation operation is performed. This involves reading a synonym table from the database, which contains multiple synonym pairs, such as (flowerpot: vase), (house: building), etc. The synonym table is then used to supplement all or some of the synonyms in the descriptive word set. For example, for the descriptive word "flowerpot," the corresponding synonym "vase" is retrieved from the synonym table and added to the descriptive word set; similarly, for the descriptive word "house," the corresponding synonym "building" is retrieved from the synonym table and added to the descriptive word set.

[0160] Next, an ambiguous word replacement operation is performed. This involves reading a replacement word table from the database, which includes multiple replacement word pairs, such as (painting: hanging painting, easel, mural), (water: bottle, lake), etc. Then, this matching replacement word table is used to replace ambiguous descriptive words in the descriptive word set. For example, for the descriptive word "water" in the descriptive word set, the corresponding replacement word "bottle" is retrieved from the replacement word table, and then "bottle" is added to the descriptive word set while "water" is removed from the descriptive word set.

[0161] Finally, an invalid word supplementation operation is performed, which involves reading the invalid word table from the database. This table contains several invalid words: sky, shadow, roof, and cloud. These invalid words are added to the descriptor set, resulting in the final descriptor set: house, building, signboard, bottle, lawn, streetlight, fence, potted plant, flowerpot, vase, wooden house, beach, tree, sky, shadow, roof, and cloud. Each descriptor in the final descriptor set serves as an entity hint for subsequent entity detection.

[0162] In this embodiment, deduplication of prompt words reduces the amount of data required for subsequent processing; error word removal avoids generating overly broad detection boxes or unidentified objects; synonym supplementation improves the accuracy of the reference image description and avoids missed detections; ambiguous word replacement reduces post-processing errors caused by ambiguous words; and invalid word supplementation avoids outputting invalid entities according to the entity categories suggested by valid words, thereby preventing the introduction of bias and errors. Through these multiple preprocessing operations, the obtained entity prompt words accurately encompass the entities of interest in the reference image, thereby improving the accuracy of subsequent entity detection.

[0163] In some embodiments, among the entity hints used for entity detection, there may be multiple entity hints corresponding to the same entity, with different segmentation granularities. When these multiple entity hints are simultaneously input into an open set detection model for entity detection, the model outputs entity candidate regions corresponding to each of these entity hints. This not only leads to duplicate detection of the target, but also, entity hints with smaller segmentation granularities will fragment the entity and output entity candidate regions of the fragmented parts of the entity, resulting in granularity ambiguity and overdetection issues. These fragmented entities will also make subsequent scene construction difficult.

[0164] For example, referring to Figure 7, the three entity prompts "potted plant," "flower," and "flowerpot" correspond to the same entity, but their segmentation granularity differs. The segmentation granularity for "potted plant" is larger, while the segmentation granularity for "flower" and "flowerpot" is smaller. When these three entity prompts are simultaneously input into an open-set detection model for entity detection, entity candidate regions 701 (for "potted plant"), 702 (for "flower"), and 703 (for "flowerpot") are detected. The entities within entity candidate regions 701 and 702 are essentially parts of the content corresponding to "potted plant" within entity candidate region 701, resulting in segmentation granularity ambiguity and over-detection issues.

[0165] To address the issues of segmentation granularity ambiguity and over-detection, this application proposes a method for entity detection that involves multiple calls to an open set detection model, specifically including the following steps:

[0166] Select multiple reference prompts that meet the preset segmentation granularity from a variety of entity prompts.

[0167] Specifically, the preset segmentation granularity is a predefined segmentation granularity adapted to the actual application scenario. The reference prompt words corresponding to the preset segmentation granularity are often entity prompt words for entities that are easily fragmented or ambiguous. For example, in a scene building task, for container objects (such as bookshelves or cabinets), it is necessary to ignore the internal objects and retain the external objects above and below. Therefore, the reference prompt word corresponding to the preset segmentation granularity is "bookshelf," not "books." Similarly, for the three entity prompt words "potted plant," "flower," and "flowerpot," the scene building task requires the entire "potted plant," not the fragmented "flower" and "flowerpot." Therefore, the reference prompt word corresponding to the preset segmentation granularity is the entity prompt word "potted plant," not the entity prompt words "flower" and "flowerpot."

[0168] Next, based on multiple reference cue words, entity detection is performed in the reference image to obtain multiple reference candidate regions. Specifically, the selected multiple reference cue words and reference image are input into an open set detection model for entity detection to obtain multiple reference candidate regions. These reference candidate regions are often detection boxes of entities that are easily segmented and fragmented, or that produce ambiguity.

[0169] Then, based on other entity cues among multiple entity cues, entity detection is performed in the reference image to obtain multiple supplementary candidate regions.

[0170] Specifically, the open-set detection model is invoked a second time to perform entity detection on the reference image based on other entity cues, thereby obtaining multiple supplementary candidate regions. In some cases, when inferring using the open-set detection model a second time, the reference image and all entity cues can be directly input into the open-set detection model for entity detection to obtain multiple supplementary candidate regions.

[0171] Among multiple supplementary candidate regions, there may be supplementary candidate regions that correspond to the same entity as the reference candidate region. Therefore, it is necessary to deduplicate the multiple supplementary candidate regions and the multiple reference candidate regions. During the deduplication process, the priority of supplementary candidate regions is lower than that of reference candidate regions. The specific deduplication process includes the following operations:

[0172] For each supplementary candidate region, the following steps are performed: if the intersection-union ratio between a supplementary candidate region and any reference candidate region is greater than the first decision threshold, then the supplementary candidate region is removed.

[0173] The first decision threshold is the criterion for judging the relationship between supplementary candidate regions and reference candidate regions. When the intersection-union ratio (IUR) between a supplementary candidate region and any reference candidate region is greater than the first decision threshold, it means that the supplementary candidate region and the reference candidate region correspond to the same entity. To avoid duplicate detection, the reference candidate region will be retained first, and the supplementary candidate region will be removed. When the IUR between a supplementary candidate region and any reference candidate region is not greater than the first decision threshold, the supplementary candidate region will be retained.

[0174] Specifically, when the intersection-union ratio (IUU) between a supplementary candidate region and any one of the multiple reference candidate regions is greater than the first decision threshold, it indicates that the supplementary candidate region and the reference candidate region correspond to the same entity. In this case, the reference candidate region is saved first, and the supplementary candidate region is removed.

[0175] Furthermore, when a supplementary candidate region is located inside any one of the multiple reference candidate regions, it indicates that the entity within the supplementary candidate region is a partial region of the entity within the reference candidate region. In this case, the reference candidate region is also saved first, and the supplementary candidate region is removed.

[0176] Finally, at least one supplementary candidate region and multiple reference candidate regions are retained as multiple entity candidate regions.

[0177] In this embodiment, entity detection is first performed on the reference image based on multiple reference prompts with a preset segmentation granularity, prioritizing the identification of reference candidate regions for entities that are easily segmented and ambiguous; then, other entity prompts are iteratively input for entity detection to obtain multiple supplementary candidate regions, and supplementary candidate regions that are duplicates of the reference candidate regions are filtered out. This avoids duplicate entity detection and also avoids generating segmented or low-quality entity candidate regions.

[0178] It should be noted that this application can also input the reference image and all entity prompts into the open set detection model for entity detection to directly obtain multiple entity candidate regions. This application does not make any specific limitations on this.

[0179] In some implementations, in the above-mentioned method of repeatedly calling the open set detection model to perform entity detection and obtain multiple entity candidate regions, the process of calling the open set detection model to perform entity detection is the same each time, the only difference being the input content. The following explanation, using the example of performing entity detection on a reference image based on multiple reference prompts to obtain multiple reference candidate regions, includes the following steps:

[0180] Feature extraction is performed on multiple reference prompt words to obtain corresponding text features; and feature extraction is performed on reference images to obtain corresponding image features; then the text features and image features are fused to obtain target fusion features; then, based on the target fusion features, multiple preliminary candidate regions in the reference images are obtained; finally, the multiple preliminary candidate regions are deduplicated to obtain multiple reference candidate regions.

[0181] In practice, multiple reference prompts are converted into low-dimensional vector representations (Embedding) using an open-set detection model to obtain the corresponding text features. The BERT model, a pre-trained language model, can be used for text feature extraction. For multiple reference prompts, they are first concatenated into a sentence, with a [CLS] marker added at the beginning and a [SEP] marker added between each reference prompt. This sentence is then input into the BERT model, which outputs the hidden state corresponding to each marker. The hidden state corresponding to the [CLS] marker is taken as the text feature of the entire sentence; this feature represents the text features of the multiple reference prompts.

[0182] In actual encoding, multiple reference prompts can be encoded separately to obtain the text encoding features of each reference prompt, and then these text encoding features can be concatenated to obtain the final text features. Alternatively, multiple reference prompts can be concatenated according to a certain format and then encoded as a whole to obtain text features. These text features can also be called text representations, text representations, or text representation vectors, etc.

[0183] The specific process of extracting features from the reference image to obtain the corresponding image features has been introduced earlier and will not be repeated here.

[0184] Since text features and image features belong to different modalities, and the feature spaces corresponding to features of different modalities are independent, text features and image features are not in the same feature space. Unaligned feature spaces will have an adverse effect on downstream tasks. Based on this, this application performs cross-modal feature fusion and alignment on the obtained text features and image features to obtain target fusion features. Specifically, cross-modal feature fusion and alignment can be performed using cross-attention mechanisms, feature weighted summation, etc., to obtain target fusion features.

[0185] Specifically, a cross-attention mechanism is used for fusion, assuming the text features are... Image features are Where m and n are the sequence lengths of text and image features, respectively, and d is the feature dimension. First, calculate the attention score S = Then, the image features are weighted and summed using attention scores to obtain the fused features F = SI, and F is used as the target fused features.

[0186] Based on the target fusion features, multiple preliminary candidate regions in the reference image are obtained through mapping. Each preliminary candidate region corresponds to an entity category and the confidence level of that entity category. Each entity category is represented by an entity cue word. Finally, the multiple preliminary candidate regions are deduplicated to obtain multiple reference candidate regions.

[0187] Wherein, it is assumed that the target fusion feature is The weight matrix of the fully connected layer is The bias vector is The mapped result is O = WF + b, where q is the output dimension. Then, based on preset thresholds and rules, multiple preliminary candidate regions are selected from O. Each preliminary candidate region corresponds to an entity category and its confidence level, and each entity category is represented by an entity cue word. The preset thresholds have different roles in different operational scenarios. When obtaining preliminary candidate regions based on target fusion feature mapping, multiple preliminary candidate regions are selected from the mapping results according to preset thresholds and rules. When deduplicating multiple preliminary candidate regions, the preset threshold serves as a stopping condition for the number of deduplication rounds; when the number of deduplication rounds reaches the preset threshold, the deduplication operation stops. When judging the area of ​​the non-overlapping region between the second pixel region and the second candidate region, if the area of ​​the non-overlapping region is greater than the preset threshold, it indicates a significant difference in range between the second pixel region and the second candidate region, and the second candidate region is a low-quality candidate region that will be directly removed. If the area of ​​the non-overlapping region is not greater than the preset threshold, the second candidate region will be retained.

[0188] In this application, entity detection is performed by combining features from the text modality and features from the image modality. The semantic information in the features of different modalities cooperates with each other, which not only maintains the advantages of each feature but also makes up for the shortcomings of a single modality feature, thereby improving the accuracy of the reference candidate regions obtained by detection.

[0189] In some implementations, the embodiments of this application employ at least the following method to deduplicat multiple preliminary candidate regions and obtain multiple reference candidate regions, specifically including the following steps: clustering multiple preliminary candidate regions according to the similarity between the entity categories corresponding to each of the multiple preliminary candidate regions to obtain multiple candidate region sets.

[0190] Specifically, each preliminary candidate region corresponds to an entity category, and each entity category is represented by an entity prompt word. Therefore, the similarity between any two entity categories can be the text similarity between the entity prompt words of the two entity categories.

[0191] Cosine similarity can be used to calculate the similarity between entity prompts for any two entity categories. Assuming the text feature vectors corresponding to two entity prompts are v1 and v2, the cosine similarity between them is... When sim(v1,v2) is greater than the similarity threshold, the corresponding preliminary candidate regions are clustered into the same candidate region set.

[0192] When the text similarity between the entity prompts of two entity categories is greater than the similarity threshold, it indicates that the entity prompts of these two entity categories are synonyms, and therefore, these two entity categories may hit the same entity. The similarity threshold can be adjusted according to the specific dataset and application scenario. For example, in common image recognition scenarios, the similarity threshold can be set between 0.7 and 0.9.

[0193] For example, referring to Figure 8A, the entity cue word "flower" corresponds to preliminary candidate region 801 and the entity cue word "lotus" corresponds to preliminary candidate region 802, both of which match the same flower. If non-maximum suppression is directly used for deduplication, the entity cue words "flower" and "lotus" will be considered to be different entity categories, thus retaining the preliminary candidate regions corresponding to these two entity categories, resulting in duplicate detection.

[0194] Based on this, this application performs synonym matching on the entity prompt words corresponding to each of the multiple preliminary candidate regions, thereby dividing the multiple preliminary candidate regions into multiple candidate region sets, each candidate region set including at least one preliminary candidate region; and then deduplicating the preliminary candidate regions in each candidate region set.

[0195] For each candidate region set, the following methods are used to remove duplicates: select the baseline candidate region with the highest confidence from the candidate region set; then remove each other preliminary candidate region in the candidate region set whose intersection-union ratio with the baseline candidate region is greater than the second decision threshold, or remove each other preliminary candidate region in the candidate region set that is located inside the baseline candidate region.

[0196] The second decision threshold is used to determine whether other preliminary candidate regions correspond to the same entity as the benchmark candidate region when deduplicating multiple preliminary candidate regions. In each candidate region set, if the intersection-union ratio (IUR) between other preliminary candidate regions and the benchmark candidate region is greater than the second decision threshold, it indicates that the two regions correspond to the same entity. Since the benchmark candidate region has the highest confidence and higher accuracy, this other preliminary candidate region will be removed. If the IUR between other preliminary candidate regions and the benchmark candidate region is not greater than the second decision threshold, but the other preliminary candidate region is located within the benchmark candidate region, it will also be removed to avoid missing or fragmented entity parts. If the IUR between other preliminary candidate regions and the benchmark candidate region is not greater than the second decision threshold, and the other preliminary candidate region is not located within the benchmark candidate region, it will be retained.

[0197] Specifically, the higher the confidence level of a preliminary candidate region, the higher its accuracy. Therefore, it is necessary to prioritize retaining preliminary candidate regions with higher confidence levels. Based on this, this application selects the preliminary candidate region with the highest confidence level from each candidate region set as the baseline candidate region; then it iterates through each other preliminary candidate region in the candidate region set to remove duplicates. For each other preliminary candidate region encountered, the intersection-union ratio (IUU) between the other preliminary candidate region and the baseline candidate region is calculated. The IUU is calculated using the following formula (1):

[0198] Where IOU represents the intersection-union ratio, A∩B represents the area of ​​the intersection region of region A and region B, and A∪B represents the area of ​​the union region of region A and region B.

[0199] If the intersection-union ratio between the other preliminary candidate region and the benchmark candidate region is greater than the second decision threshold, it means that the other preliminary candidate region and the benchmark candidate region correspond to the same entity. Therefore, the other preliminary candidate region with lower confidence is removed first.

[0200] For example, regarding the preliminary candidate region 801 corresponding to the entity prompt word "flower" and the preliminary candidate region 802 corresponding to the entity prompt word "lotus" shown in Figure 8A, since the text similarity between the entity prompt words "flower" and "lotus" is greater than a preset threshold, preliminary candidate regions 801 and 802 will be assigned to the same candidate region set for deduplication. Then, during deduplication, if the intersection-union ratio (IUU) between preliminary candidate regions 801 and 802 is greater than the second decision threshold, and the confidence level of preliminary candidate region 801 corresponding to the entity prompt word "flower" is lower, then preliminary candidate region 801 will be removed.

[0201] If the intersection-union ratio (IUU) between the other preliminary candidate region and the baseline candidate region is not greater than the second decision threshold, it is possible that the baseline candidate region may contain the other preliminary candidate region. That is, the entities in the other preliminary candidate region may be partial regions of the entities in the baseline candidate region. Based on this, to avoid missing or fragmented entities, the other preliminary candidate region is also removed.

[0202] For example, referring to Figure 8B, the preliminary candidate region 803 corresponding to the entity prompt word "bookshelf" surrounds the preliminary candidate region 804 corresponding to the entity prompt word "book". Therefore, the preliminary candidate region 804 corresponding to "book" is removed.

[0203] It should be noted that after deduplicating each of the other preliminary candidate regions in the candidate region set, if the candidate region set still retains multiple other preliminary candidate regions, the same method as before is used to deduplicatize the remaining multiple preliminary candidate regions again, i.e., a new round of deduplication is started, and so on, until a preset stopping condition is met. The stopping condition can be that the number of deduplication rounds reaches a preset threshold, or that no other preliminary candidate regions besides the baseline candidate region are retained after deduplication. The setting of the threshold affects the effectiveness and efficiency of deduplication. If the preset threshold is set too low, it may stop before deduplication is sufficient, resulting in some duplicate candidate regions remaining; if the preset threshold is set too high, it will increase the computational load and time cost of deduplication. In practical applications, the preset threshold can be adjusted according to the characteristics of the dataset and the requirements of the task. For example, for datasets with a large number of entities and complex duplication situations, the preset threshold can be appropriately increased; for datasets with a small number of entities and simple duplication situations, the preset threshold can be appropriately decreased.

[0204] Finally, based on at least one preliminary candidate region retained by each of the multiple candidate region sets, multiple reference candidate regions are obtained.

[0205] In this embodiment, multiple preliminary candidate regions are clustered according to the similarity between the entity categories corresponding to each preliminary candidate region, resulting in multiple candidate region sets. Then, deduplication is performed separately for each candidate region set. This effectively detects cases where similar categories correspond to the same entity, thereby improving the accuracy of deduplication and avoiding duplicate detection. Secondly, during the deduplication process, the enclosing relationship between candidate regions is considered, and candidate regions enclosing the outer regions are retained, thereby avoiding duplication, partial omission, or fragmentation in entity detection.

[0206] Step 402: Select at least one first candidate region and at least one second candidate region from multiple entity candidate regions. The confidence level of the first candidate region is greater than the first threshold, and the confidence level of the second candidate region is not greater than the first threshold.

[0207] Specifically, after selecting at least one first candidate region with a confidence level greater than a first threshold from multiple entity candidate regions, all entity candidate regions with a confidence level not greater than the first threshold can be retained for subsequent supplementation of missed entities; alternatively, a second threshold can be set, wherein the second threshold is less than the first threshold. At least one second candidate region with a confidence level greater than the second threshold and a confidence level not greater than the first threshold is then selected from multiple entity candidate regions, and the selected second candidate region is used for subsequent supplementation of missed entities.

[0208] For example, if the first threshold is set to 0.5 and the second threshold is set to 0.3, firstly select the first candidate region (i.e., the high-confidence candidate region) with a confidence level greater than 0.5 from multiple entity candidate regions; then retain the second candidate region (i.e., the low-confidence candidate region) with a confidence level greater than 0.3 and not greater than 0.5 from the remaining entity candidate regions.

[0209] Step 403: Segment the first pixel region containing the entity from each first candidate region; and segment the second pixel region containing the entity from each second candidate region.

[0210] Specifically, for each first candidate region, the first candidate region is input into the image segmentation model for inference. In the image segmentation model, features are extracted from the first candidate region to obtain corresponding image features. Based on the obtained image features, the class label of each pixel in the first candidate region is predicted. Multiple pixels with class labels corresponding to the entity class of the first candidate region are concatenated to obtain the first pixel region; this first pixel region is also the pixel mask of the entity within the first candidate region. Assume the image segmentation model is a fully convolutional network (FCN) based on a convolutional neural network (CNN). Image features are input to the decoding part of the model, and after a series of deconvolutional and convolutional layers, the size of the feature map is gradually restored to be the same as the input first candidate region. Each pixel position corresponds to an output vector, and the dimension of the vector is equal to the number of entity classes. The softmax function is used to convert the output vector into a probability distribution for each entity class, and the class with the highest probability is selected as the class label of that pixel.

[0211] Image feature extraction can be performed using convolutional neural networks (CNNs). Assuming the first candidate region is an image of size h1×w1×c1 (h1 and w1 are the height and width, respectively, and c1 is the number of channels), it can be processed through a series of convolutional layers, pooling layers, and activation functions. For example, after a convolutional layer with a kernel size of k×k, a stride of s, padding of p, weights of W, and a bias of b, the output Y of the convolutional layer can be expressed as Y = ReLU(conv(X,W) + b), where conv represents the convolution operation and ReLU is the activation function. After processing through multiple such layers, the image features are obtained.

[0212] In some embodiments, the first candidate region can be expanded outward according to a preset ratio to obtain an expanded candidate region; then, the expanded candidate region is input into an image segmentation model for entity segmentation to obtain the first pixel region, thereby improving the image segmentation effect. Assuming the coordinates of the first candidate region are (x1, y1, x2, y2) (the top-left corner coordinates are (x1, y1), and the bottom-right corner coordinates are (x2, y2)), and the preset ratio is r, then the top-left corner coordinates of the expanded candidate region are (x1-r·(x2-x1), y1-r·(y2-y1)), and the bottom-right corner coordinates are (x2+r·(x2-x1), y2+r·(y2-y1)), while ensuring that the expanded coordinates are within the image range. The preset ratio is the proportional standard used when expanding the first or second candidate region. When performing entity segmentation, the candidate region is expanded outward according to a preset ratio to obtain an expanded candidate region. The expanded candidate region is then input into the image segmentation model for entity segmentation. The purpose of this is to improve the image segmentation effect. If the expansion operation is not performed according to the preset ratio, more accurate segmentation results may not be obtained.

[0213] The second pixel region of each second candidate region can be obtained using the same method described above, which will not be repeated here.

[0214] Step 404: Select at least one target pixel region from the obtained at least one second pixel region; wherein the proportion of each target pixel region in the second candidate region from which the target pixel region is segmented is greater than a first proportion threshold.

[0215] Specifically, the target pixel region is the pixel region in each second pixel region that meets the quality verification conditions, wherein the quality verification conditions include at least: the proportion of the second pixel region in the corresponding second candidate region is greater than the first proportion threshold.

[0216] The first proportion threshold is a key indicator used to filter target pixel regions from at least one second pixel region. When the proportion of a second pixel region in its corresponding second candidate region is greater than the first proportion threshold, the second pixel region is eligible to be considered as a target pixel region for subsequent supplementation of missed entities. When the proportion of a second pixel region in its corresponding second candidate region is not greater than the first proportion threshold, it indicates that the second pixel region may be a thin, elongated object or an object that cannot be correctly segmented. In this case, the second pixel region will be excluded and will not be used for subsequent supplementation of missed entities.

[0217] In practical applications, when the proportion of the second pixel region in the corresponding second candidate region is too low or the central region is empty, it indicates that the second pixel region may be a thin and elongated small object or an object that cannot be correctly segmented. Therefore, it is not suitable for subsequent supplementation of missed entities and is directly removed.

[0218] In some embodiments, the quality verification condition, in addition to the proportion of the second pixel region in the corresponding second candidate region being greater than the first proportion threshold, also includes at least one of the following:

[0219] The size of the second candidate region is within a preset range;

[0220] The confidence level corresponding to the second candidate region is greater than the second threshold;

[0221] The preset interval is a standard range used during the quality verification process to determine whether the size of the second candidate region is appropriate. When the size of the second candidate region (which can be characterized by parameters such as area) is within the preset interval, the second candidate region is eligible to meet the quality verification conditions and can be used to supplement subsequently missed entities; when the size of the second candidate region is not within the preset interval, for example, if the area is too large, it may contain errors, while if the area is too small, it is usually not important, and in this case, the second candidate region will be removed.

[0222] Specifically, the size of the second candidate region can be characterized by its area or by other parameters. If the area of ​​the second candidate region is too large, it may contain errors; if the area of ​​the second candidate region is too small, it is usually unimportant and can therefore be removed. That is, this application retains second candidate regions with areas within a preset range for subsequent supplementation of missed entities.

[0223] Secondly, a second threshold is set, which is less than the first threshold. If the confidence level of the second candidate region is less than the second threshold, it indicates an error and can therefore be removed. In other words, this application retains the corresponding second candidate regions with a confidence level greater than the second threshold for subsequent supplementation of missed entities.

[0224] It should be noted that if the previous steps used the condition that the confidence level is greater than the second threshold and the confidence level is not greater than the first threshold to select at least one second candidate region from multiple entity candidate regions, then when selecting the target pixel region, the quality check condition that the confidence level of the second candidate region is greater than the second threshold can be removed to avoid duplicate processing.

[0225] In some embodiments, if, during the process of obtaining the second pixel region, the second candidate region is first expanded outward according to a preset ratio to obtain an expanded candidate region; and then the expanded candidate region is input into an image segmentation model for entity segmentation to obtain the second pixel region, then the quality verification conditions further include:

[0226] The area of ​​the non-overlapping region between the second pixel region and the second candidate region is no greater than a preset threshold.

[0227] Specifically, when the area of ​​the non-overlapping region between the second pixel region and the second candidate region is greater than a preset threshold, it indicates that the range of the second pixel region and the range of the second candidate region are significantly different, meaning that the second pixel region and the second candidate region do not correspond. Therefore, the entity in the second candidate region is likely not a complete entity, meaning that the second candidate region is a low-quality candidate region, and thus it can be directly removed.

[0228] In this embodiment, the quality of the selected second candidate region and the corresponding second pixel region is checked, and the second candidate region with low quality or containing incomplete entities and the corresponding second pixel region are filtered out, thereby improving the effect of subsequent supplementation of missed entities.

[0229] Step 405: Obtain the segmentation result of the reference image based on at least one first pixel region and at least one target pixel region.

[0230] Specifically, each first pixel region corresponds to an entity candidate region and the entity category corresponding to that entity candidate region; similarly, each target pixel region corresponds to an entity candidate region and the entity category corresponding to that entity candidate region.

[0231] The obtained first pixel regions, corresponding entity candidate regions and entity categories, as well as each target pixel region, corresponding entity candidate regions and entity categories, can be directly used as the segmentation results of the reference image; alternatively, a series of post-processing steps can be performed on the obtained first pixel regions and each target pixel region to obtain the segmentation results of the reference image. This application does not specifically limit the specific application in this regard.

[0232] In this embodiment, entity categories associated with a reference image are used as prompts to perform entity detection in the reference image, obtaining multiple entity candidate regions. The entity categories associated with the reference image are not limited to those labeled during the training phase, but may also include entity categories not labeled during the training phase. Therefore, when performing entity detection, entity candidate regions corresponding to entity categories labeled during the training phase and entity candidate regions corresponding to entity categories not labeled can be detected, thereby enabling the multiple entity candidate regions obtained to cover all entities to be identified as much as possible, improving the comprehensiveness of entity detection.

[0233] Secondly, from multiple entity candidate regions, a first candidate region with a confidence level greater than a first threshold (i.e., a high-confidence candidate region) is selected, while a second candidate region with a confidence level not greater than the first threshold (i.e., a low-confidence candidate region) is retained. Then, each first candidate region is segmented into pixels to obtain a corresponding first pixel region, thus identifying the pixel regions of salient entities in the reference image. Simultaneously, each second candidate region is segmented into pixels to obtain a corresponding second pixel region, and a target pixel region with a larger proportion in each second candidate region is selected to identify the pixel regions of missed entities. Finally, based on the obtained first pixel regions and target pixel regions, the segmentation result of the reference image is generated. This not only identifies salient entities in the reference image but also uses low-confidence candidate regions to supplement undetected entities, greatly reducing the number of missed entities and thus improving the accuracy of entity recognition in the image.

[0234] In addition, in the scene building task, low-confidence entity candidate regions are retained and used to supplement entities that were not successfully detected. In this way, even if there are differences between the image style and labeled entity categories of the dataset used during training and the actual image style and expected entity categories of the reference images in the scene building task, entity omission problems can be detected and corrected in time, thereby improving the performance of the scene building task.

[0235] In some embodiments, incremental verification is performed on at least one first pixel region and at least one target pixel region to obtain a segmentation result of the reference image. Specifically, this includes the following steps: constructing an initial set of confidence boxes based on the first candidate regions of each of the at least one first pixel region; and stitching the at least one first pixel region together to form an initial stitched pixel region. Based on the coverage relationship between the second candidate regions of each of the at least one target pixel region and the confidence box set, the confidence box set and the stitched pixel region are iteratively updated to obtain a segmentation result of the reference image. Each iterative update process includes the following operations: if a second candidate region covers multiple candidate regions in the previously updated confidence box set, the second candidate region is added to the previously updated confidence box set, and multiple candidate regions are removed from the previously updated confidence box set to obtain the currently updated confidence box set; the target pixel region corresponding to the second candidate region is stitched together to the previously updated stitched pixel region to obtain the currently updated stitched pixel region.

[0236] Specifically, firstly, based on the first candidate regions corresponding to at least one first pixel region, an initial set of confidence boxes is constructed, and the at least one first pixel region is concatenated into an initial concatenated pixel region. Next, for each target pixel region, a second candidate region is iteratively updated according to its coverage relationship with the candidate regions in the current set of confidence boxes. In each iteration, if a second candidate region covers multiple candidate regions in the current set of confidence boxes, the second candidate region is added to the set of confidence boxes, the covered candidate regions are removed, and the target pixel region corresponding to the second candidate region is concatenated into the current concatenated pixel region. After traversing and processing all second candidate regions, the final reference image segmentation result is obtained.

[0237] If a second candidate region covers multiple candidate regions in the confidence box set, it means that the entities in the multiple candidate regions are entity components or internal entities of a complete entity, while the entities in the second candidate region are a complete entity.

[0238] To improve the accuracy of recognition, the connectivity between the pixel regions segmented from each of the multiple candidate regions can be further determined, i.e., whether a continuous pixel region is formed. If so, it indicates that there is a complete entity within the second candidate region. Therefore, the multiple candidate regions in the confidence box set are replaced with the second candidate region. At the same time, the pixel regions corresponding to the multiple candidate regions in the spliced ​​pixel region are replaced with the target pixel region segmented from the second candidate region.

[0239] After traversing all the second candidate regions, the iteration to update the confidence box set and the stitched pixel region stops, resulting in the final confidence box set and the final stitched pixel region. The candidate regions in the final confidence box set, along with their corresponding entity categories, and the pixel regions corresponding to each entity category in the final stitched pixel region, constitute the segmentation result of the reference image.

[0240] The process of determining the connectivity between the pixel regions segmented from multiple candidate regions—that is, whether they form a continuous pixel region—can be achieved using a breadth-first search (BFS) algorithm. First, the pixel regions segmented from each candidate region are merged into a pixel set P. Then, a pixel is randomly selected from set P as a starting point s, marked as visited, and added to queue Q. While the queue is not empty, the pixel p at the head of the queue is retrieved, and its adjacent pixels n are traversed. If n belongs to set P and has not been visited, it is marked as visited and added to queue Q. When the queue is empty, it is checked whether all pixels belonging to set P have been visited. If so, the pixel regions segmented from the multiple candidate regions are connected, forming a continuous pixel region; otherwise, they are not connected. If they are connected, it indicates that the second candidate region contains a complete entity. Therefore, the multiple candidate regions in the confidence box set are replaced with the second candidate region; simultaneously, the pixel regions corresponding to the multiple candidate regions in the concatenated pixel region are replaced with the target pixel region segmented from the second candidate region.

[0241] In some embodiments, before replacing multiple candidate regions in the confidence box set with the second candidate region, the target pixel region segmented from the second candidate region can be quality-assessed. Specific assessment methods include, but are not limited to, detection box range assessment, geometric structure assessment, and semantic similarity assessment. If the quality of the target pixel region is unsatisfactory, the second candidate region can be re-segmented to obtain a new target pixel region. Then, the multiple candidate regions in the confidence box set can be replaced with the second candidate region, and the pixel regions corresponding to each of the multiple candidate regions in the stitched pixel region can be replaced with the new target pixel region segmented from the second candidate region.

[0242] In this embodiment, multiple small candidate regions with high confidence are replaced with a large candidate region containing a complete entity. This effectively avoids fragmented entity segmentation, thereby improving the accuracy of target detection and segmentation. Secondly, by performing novelty checks on the retained low-confidence candidate regions, the candidate regions for vulnerable entities can be effectively supplemented, thereby reducing the number of missed entities.

[0243] In some embodiments, during each iteration update process, if a second candidate region does not contain multiple candidate regions, the target pixel region corresponding to the second candidate region is determined as the newly added region compared to the previously updated stitched pixel region.

[0244] When the proportion of the newly added region in the target pixel region is greater than the second proportion threshold, the second candidate region is added to the previously updated confidence box set to obtain the current updated confidence box set; and the target pixel region is stitched to the previously updated stitched pixel region to obtain the current updated stitched pixel region.

[0245] The second proportional threshold is a crucial criterion for determining whether a target pixel region corresponds to a newly added entity during the novelty verification process. Specifically, when the proportion of the newly added region within the target pixel region exceeds the second proportional threshold, it indicates that the target pixel region corresponds to a newly added entity. Therefore, the second candidate region containing the target pixel region is added as a new bounding box to the confidence box set. Simultaneously, the target pixel region is concatenated to the concatenated pixel region to perform segmentation mask correction on the concatenated pixel region. When the proportion of the newly added region within the target pixel region is not greater than the second proportional threshold, it indicates that the target pixel region may be a component or internal entity of an entity within a candidate region of the confidence box set. In this case, the second candidate region containing the target pixel region is directly removed, and then the target pixel region is concatenated to the concatenated pixel region to perform segmentation mask correction on the concatenated pixel region.

[0246] Similarly, by adding the second candidate region as a new bounding box before the set of confidence boxes, the quality of the target pixel region segmented from the second candidate region can also be judged. The specific judgment process has been introduced above and will not be repeated here.

[0247] For example, referring to Figure 9A, multiple first candidate regions (i.e., high-confidence candidate regions) with confidence scores greater than a first threshold are selected from the multiple entity candidate regions output by the open set detection model. An initial set of confidence boxes is constructed based on these high-confidence detection boxes. A full segmentation model is used to perform foreground segmentation on each high-confidence detection box to obtain the corresponding segmentation mask. The obtained multiple segmentation masks are then concatenated to form an initial concatenated pixel region.

[0248] Multiple second candidate regions (i.e., low-confidence candidate regions) with a confidence level greater than a second threshold but not greater than a first threshold are selected from multiple entity candidate regions. A full segmentation model is then used to perform foreground segmentation on each low-confidence candidate region to obtain the corresponding segmentation mask.

[0249] First, a quality check is performed on each low-confidence candidate region. The quality check conditions include: 1. The area of ​​the low-confidence candidate region is within a preset range, i.e., the area is moderate; 2. The proportion of the segmentation mask in the corresponding low-confidence candidate region is greater than a first proportion threshold, i.e., the foreground proportion is high; 3. The area of ​​the non-overlapping region between the low-confidence candidate region and the corresponding segmentation mask is not greater than a preset threshold, i.e., the segmentation and detection correspond. When a low-confidence candidate region meets the above three quality check conditions, it is determined that the quality check of the low-confidence candidate region has passed.

[0250] For each low-confidence candidate region that passes the quality check, an augmentation check is performed. Specifically, if a low-confidence candidate region covers multiple candidate regions in the confidence box set, the low-confidence candidate region is added to the confidence box set as a new box, and the multiple covered candidate regions in the confidence box set are removed. Secondly, the segmentation mask in the low-confidence candidate region is used to correct the segmentation mask of the stitched pixel region.

[0251] If a low-confidence candidate region does not cover multiple candidate regions in the confidence box set, then the segmentation mask in the low-confidence candidate region is determined to be an additional region compared to the stitched pixel region.

[0252] If the proportion of the newly added region in the segmentation mask of the low-confidence candidate region is greater than the second proportion threshold (i.e., the area of ​​the newly added region is large enough), then the low-confidence candidate region is added as a new bounding box to the set of confidence boxes; then the segmentation mask of the low-confidence candidate region is used to correct the segmentation mask of the stitched pixel region.

[0253] If the proportion of the newly added region in the segmentation mask of the low-confidence candidate region is not greater than the second proportion threshold, the low-confidence candidate region is discarded, and the segmentation mask of the low-confidence candidate region is used to correct the segmentation mask of the stitched pixel region.

[0254] In this embodiment, based on the segmentation mask in the low-confidence candidate region, the missed entities in the reference image are identified compared to the area of ​​the newly added region in the stitched pixel region. Then, the missed entities are added to the segmentation result in the reference image, thereby reducing the number of missed entities and improving the accuracy of entity recognition in the image.

[0255] In some embodiments, since invalid words are included in the multiple entity prompts input during entity detection, the segmentation results of the reference image may contain segmentation results corresponding to invalid words. Therefore, the segmentation results corresponding to each invalid word are removed from the segmentation results of the reference image to ensure that subsequent processing is based only on the segmentation results corresponding to valid words.

[0256] In some embodiments, scene building tasks typically involve three types of entities: main bodies, decorative elements, and background entities. Main bodies (such as tables, chairs, and beds) generally have a larger area and a greater impact on scene building; their absence or errors will significantly affect the final result. Decorative elements (such as small vases on tables) generally have a smaller area and a relatively smaller impact on scene building, and can be filled using other methods. Background entities refer to entities that affect the overall layout (such as floors, doors, windows, lawns, and water surfaces), providing the basic structure for the scene. The importance of different types of entities varies, and different types of entities may also require different post-processing steps. For example, main bodies, decorative elements, and background entities can be processed by an entity recommendation system, a decoration enhancement system, and a terrain generation system, respectively, to build the scene.

[0257] However, open set detection models and image segmentation models do not consider downstream applications, and therefore do not distinguish between subjects, decorations, and scene entities, which makes subsequent scene construction tasks difficult.

[0258] Based on this, in this application, the segmentation result of the reference image includes: segmentation results of multiple entities, and the segmentation result of each entity includes: a detection box, the entity category corresponding to the detection box, and the pixel region segmented from the detection box. According to the entity categories corresponding to the multiple entities, the scene entity is first separated from the multiple entities; then, according to the pixel region area, the remaining entities are divided into main bodies and decorative bodies. For example, pixel regions with an area greater than a preset threshold are considered main bodies, and pixel regions with an area not greater than the preset threshold are considered decorative bodies. The segmentation results of the main body, the decorative body, and the scene entity are then input into subsequent systems for use, thereby improving the accuracy and efficiency of subsequent processing.

[0259] In some embodiments, during scene building tasks, entity detection aims to detect as many entities as possible. To avoid missing entities, open set detection models and image segmentation models independently distinguish as many entities as possible. This results in segmentation results for the reference image containing a large number of small entities (such as small items on shelves or small decorations on doors and windows). The segmentation results of these small entities have little impact on scene building, but may lead to poor visualization and increase the processing time and cost of downstream tasks.

[0260] Therefore, this application filters out an excessive number of small entities. The specific filtering scheme is as follows: a maximum number of entities is set for the reference image; when the number of candidate regions detected in the reference image is greater than the maximum number of entities set, the segmentation results of entities whose pixel area is lower than a set threshold are filtered out to ensure that the remaining entities have certain visual and functional value.

[0261] The maximum number of entities is an upper limit set for the reference image in the scene building task. When the number of candidate regions detected in the reference image is greater than the maximum number of entities, it means that there may be too many small entities that have little impact on scene building. In this case, the segmentation results of entities with pixel area below the set threshold will be filtered out to ensure that the remaining entities have certain visual and functional value. When the number of candidate regions detected in the reference image is not greater than the maximum number of entities, no filtering operation is required.

[0262] Identifying small entities can be achieved by calculating the pixel area of ​​each entity. For each entity's segmentation result, the number of pixels within its pixel area is counted as the area. An area threshold is set; when an entity's pixel area is less than this threshold, it is considered a small entity. The filtering algorithm is as follows: First, all entities are sorted from largest to smallest according to their pixel area. Then, the number of entities is counted. If the number exceeds the maximum number of entities, the pixel area of ​​each entity in the sorted entity list is checked sequentially, and entities with areas below the set threshold are filtered out until the number of entities does not exceed the maximum number of entities.

[0263] In some cases, the maximum number of entities can be set separately for the main body, decorative entities, and scene entities, thereby avoiding the detection of too many entities.

[0264] In addition, to ensure the richness of the constructed scene, it is also necessary to avoid the number of entities used to construct the scene being too low. Therefore, a minimum number of entities is set for the reference image. In some cases, the minimum number of entities can also be set for the main body, decorations, and scenery entities respectively, thereby ensuring the richness of the scene.

[0265] The minimum number of entities is a lower limit set for the number of reference images to ensure the richness of the scene construction. In some cases, this number may also be set separately for main objects, decorative objects, and background objects. When the final number of entities determined in the reference image is lower than the minimum number of entities, the constructed scene may not be rich enough, and further adjustments to the detection or filtering strategy are needed; when the final number of entities determined in the reference image is not lower than the minimum number of entities, it indicates that the richness of the scene is guaranteed.

[0266] In some embodiments, different entities have varying degrees of importance in scene building tasks. For example, large entities like the main subject are more important for scene building, while smaller entities like decorations have less impact. To ensure the smooth progress of downstream scene building tasks, it is necessary to sort the other entities (i.e., the main subject and decorations) according to their importance, including at least the following sorting methods:

[0267] Method 1: Sort by pixel area. This involves sorting multiple entities from largest to smallest pixel area to obtain the sorted result. In scene building tasks, entities ranked higher are more important and should be placed first. Additionally, if a lower-ranked entity is a decoration; for example, an entity ranked after Nth is a decoration, then the decoration after Nth can be separated and the segmentation result is output to the downstream scene enhancement system for further processing to improve the overall aesthetics of the scene. Here, N is a positive integer.

[0268] The scene is built using the segmentation results of the retained N entities and the segmentation results of each scene entity. During the actual construction process, the materials of the N entities are placed sequentially according to the sorting result of the N entities, thereby improving the scene construction effect.

[0269] Method two involves sorting by placement relationship; that is, first identifying the placement relationship between entities. For example, if entity A is placed on entity B, then entity B must be placed first, followed by entity A. Then, based on the placement relationship between entity categories, algorithms such as topological sorting are used to sort the entities to obtain the sorted result. During scene construction, the assets of each entity are placed sequentially according to the sorted result.

[0270] The placement relationship can be identified by combining the results of object detection and semantic segmentation with spatial location information. First, object detection is used to obtain the bounding boxes of each entity, and semantic segmentation is used to obtain the pixel masks of each entity. For two entities A and B, the center coordinates (x, y, y) of their bounding boxes are calculated. A ,y A ) and (x B ,y B ) and height h A and h B If y A <y B +h B And y A >y B If the pixel masks of the two entities overlap, it can be preliminarily determined that entity A is placed on entity B.

[0271] The graph structure required for topological sorting can be constructed through the following steps: First, treat each entity as a node in the graph. Then, add directed edges to the graph based on the identified placement relationships. If entity A is placed on entity B, then add a directed edge from node B to node A, indicating that B needs to be placed first. By traversing all placement relationships, a complete directed graph is constructed. During the graph construction process, it is necessary to avoid adding edges repeatedly and to handle potential circular dependencies. If circular dependencies occur, adjustments can be made according to the actual situation, such as breaking the cycle based on the importance of the entities or other rules.

[0272] Additionally, if the upstream system or the object being operated inputs an entity prompt word, and the segmentation result in the reference image contains the segmentation result corresponding to that entity prompt word, then it is necessary to ensure that the entity corresponding to that entity prompt word is displayed first. That is, the entity should be placed at the top first, and then other entities should be sorted according to method one or method two mentioned above, so as to ensure the effect of scene construction.

[0273] To better explain the embodiments of this application, the following describes an image recognition method provided by the embodiments of this application in conjunction with a specific implementation scenario. The process of this method can be executed by a computer device and includes the following steps, as shown in Figure 9B:

[0274] The system acquires a reference image for scene construction and receives the topic descriptor "seaside cabin" from the upstream input of the reference image. It then selects a hot word library that matches the scene category of the reference image. The reference image, the topic descriptor, and multiple supplementary descriptors from the hot word library are input into the prompt word acquisition unit.

[0275] In the cue word acquisition unit, the topic descriptor "seaside cabin" is input into the cue word expansion model to obtain multiple target descriptors for the reference image; the reference image is then input into the image description model to obtain multiple preliminary descriptors for the reference image. Preprocessing is performed on the descriptor set consisting of multiple target descriptors, multiple preliminary descriptors, and multiple supplementary descriptors to obtain multiple entity cue words, namely: chair, building, fence, bottle, lawn, etc. The reference image and multiple entity cue words are then input into the entity detection unit.

[0276] In the entity detection unit, multiple entity prompts are divided into: multiple reference prompts and other entity prompts. First, an open-set detection model is used to perform entity detection on the reference image based on multiple reference prompts with a preset segmentation granularity, obtaining reference candidate regions for entities that are easily fragmented or ambiguous. Then, the open-set detection model is called a second time to perform entity detection on the reference image based on other entity prompts, obtaining multiple supplementary candidate regions. Supplementary candidate regions that overlap with the reference candidate regions are filtered out, resulting in the final multiple entity candidate regions.

[0277] First, select high-confidence candidate regions with a confidence level greater than 0.5 from multiple entity candidate regions. These are: entity candidate regions corresponding to the entity prompt word "chair" (confidence level of 82%), entity candidate regions corresponding to the entity prompt word "fence" (confidence level of 74%), entity candidate regions corresponding to the entity prompt word "food" (confidence level of 72%), etc.

[0278] Then, candidate regions with low confidence levels greater than 0.3 and not greater than 0.5 are retained, namely: the candidate region corresponding to the entity cue word "road sign" (confidence level of 46%) and the candidate region corresponding to the entity cue word "roof" (confidence level of 0.4).

[0279] Multiple high-confidence candidate regions and multiple low-confidence candidate regions are input into the entity segmentation and post-processing unit. In the entity segmentation and post-processing unit, foreground segmentation is performed on each high-confidence candidate region using an image segmentation model to obtain the corresponding pixel mask; and foreground segmentation is performed on each low-confidence candidate region using the same image segmentation model to obtain the corresponding pixel mask.

[0280] For each low-confidence candidate region, a quality check is performed. If the quality check passes, an augmentation check is performed on each low-confidence candidate region by combining each high-confidence candidate region and its corresponding pixel mask. After the augmentation check is completed, the segmentation results of multiple entities are output to the entity sorting and separation unit.

[0281] In the entity sorting and separation unit, the scene entities are first separated from multiple entities. Then, the other entities are sorted according to their pixel mask areas from largest to smallest to obtain the sorting results. The segmentation results of the top N entities in the sorting results, along with the segmentation results of each scene entity, are input into the downstream system for scene construction. The segmentation results include: bounding boxes, entity categories, and pixel masks.

[0282] In this embodiment, the reference image for scene construction has been effectively adapted to the application scenario. By combining open set detection, full segmentation, and post-processing strategies, it can better adapt to UGC scene construction tasks. Based on industry-leading open set entity full recognition and segmentation, this solution effectively alleviates the problem of entity granularity ambiguity and reduces issues such as duplicate object detection and missed detection. This provides a more effective solution for entity understanding of images or integration into downstream construction systems.

[0283] Secondly, this application has lower requirements for training data and data domain, and can still achieve good results with less labeled data and diverse input conditions. This flexibility makes the system more operable and adaptable in practical applications.

[0284] Furthermore, the strategy in this application offers good controllability, allowing for adjustments and expansion based on specific needs. This controllability not only enhances the system's flexibility but also facilitates future functional expansion and optimization. The target objects can be flexibly adjusted to achieve optimal results based on different application scenarios and requirements.

[0285] Based on the same technical concept, this application provides a schematic diagram of the structure of an image recognition device, as shown in Figure 10. The device 1000 includes:

[0286] The detection module 1001 is used to perform entity detection in the reference image based on the entity categories associated with the reference image, and obtain multiple entity candidate regions. Each entity candidate region corresponds to an entity category and a confidence level of that entity category. The confidence level represents the probability that an entity in the corresponding entity candidate region belongs to the entity category.

[0287] The filtering module 1002 is used to filter at least one first candidate region and at least one second candidate region from the plurality of entity candidate regions, wherein the confidence level of the first candidate region is greater than a first threshold and the confidence level of the second candidate region is not greater than the first threshold.

[0288] The segmentation module 1003 is configured to segment a first pixel region containing an entity from each first candidate region; and to segment a second pixel region containing an entity from each second candidate region.

[0289] The filtering module 1002 is further configured to filter at least one target pixel region from the obtained at least one second pixel region; wherein the proportion of each target pixel region in the corresponding second candidate region is greater than a first proportion threshold.

[0290] The recognition module 1004 is used to obtain the segmentation result of the reference image based on the at least one first pixel region and the at least one target pixel region.

[0291] Optionally, the detection module 1001 is further configured to:

[0292] Feature extraction is performed on the reference image to obtain the corresponding image features;

[0293] Based on the image features, the scene category and multiple preliminary descriptive words of the reference image are obtained; each preliminary descriptive word is used to describe an entity initially identified from the reference image.

[0294] From multiple preset word libraries, a target word library matching the scene category is selected, wherein the target word library includes multiple supplementary descriptive words; each supplementary descriptive word is used to describe an entity that occurs more frequently than a preset threshold under the scene category;

[0295] Based on the multiple preliminary descriptive words and the multiple supplementary descriptive words, multiple entity prompt words are obtained; wherein, each entity prompt word represents an entity category.

[0296] Optionally, the detection module 1001 is specifically used for:

[0297] Obtain the topic descriptive words corresponding to the reference image settings;

[0298] Based on the semantic information contained in the topic descriptors, the topic descriptors are expanded to obtain multiple target descriptors;

[0299] Based on the multiple target descriptors, the multiple preliminary descriptors, and the multiple supplementary descriptors, multiple entity prompt words are obtained.

[0300] Optionally, the detection module 1001 is specifically used for:

[0301] Based on the multiple target descriptors, the multiple preliminary descriptors, and the multiple supplementary descriptors, a descriptor set is constructed;

[0302] Perform a preprocessing operation on the descriptor set, and use multiple descriptors in the preprocessed descriptor set as the multiple entity prompt words;

[0303] The preprocessing operations include at least one of the following: deduplication of prompt words; removal of erroneous words; supplementation of synonyms; replacement of ambiguous words; and supplementation of invalid words.

[0304] Optionally, the detection module 1001 is specifically used for:

[0305] Select multiple reference prompt words that meet the preset segmentation granularity from the multiple entity prompt words;

[0306] Based on the plurality of reference cue words, entity detection is performed in the reference image to obtain a plurality of reference candidate regions; and based on other entity cue words among the plurality of entity cue words, entity detection is performed in the reference image to obtain a plurality of supplementary candidate regions;

[0307] For each supplementary candidate region, the following steps are performed: if the intersection-union ratio between a supplementary candidate region and any of the reference candidate regions is greater than the first decision threshold, then the supplementary candidate region is removed.

[0308] At least one supplementary candidate region and the plurality of reference candidate regions are retained as the plurality of entity candidate regions.

[0309] Optionally, the detection module 1001 is specifically used for:

[0310] Feature extraction is performed on the plurality of reference prompt words to obtain corresponding text features; and feature extraction is performed on the reference image to obtain corresponding image features;

[0311] The text features and the image features are fused to obtain the target fused features;

[0312] Based on the target fusion features, multiple preliminary candidate regions in the reference image are obtained;

[0313] The multiple preliminary candidate regions are deduplicated to obtain the multiple reference candidate regions.

[0314] Optionally, the detection module 1001 is specifically used for:

[0315] Based on the similarity between the entity categories corresponding to each of the multiple preliminary candidate regions, the multiple preliminary candidate regions are clustered to obtain multiple candidate region sets;

[0316] For each candidate region set, the following steps are performed: selecting a baseline candidate region with the highest confidence from the candidate region set; removing each other preliminary candidate region in the candidate region set whose intersection-union ratio with the baseline candidate region is greater than the second decision threshold, or removing each other preliminary candidate region in the candidate region set that is located inside the baseline candidate region;

[0317] Based on at least one preliminary candidate region retained by each of the multiple candidate region sets, the multiple reference candidate regions are obtained.

[0318] Optionally, the filtering module 1002 is specifically used for:

[0319] For each second pixel region, perform the following:

[0320] If the proportion of a second pixel region in the corresponding second candidate region is greater than the first proportion threshold, and the size of the second candidate region is within a preset range, and the confidence level of the second candidate region is greater than the second threshold, then the second pixel region is determined to be the target pixel region.

[0321] Optionally, the identification module 1004 is specifically used for:

[0322] Based on the first candidate regions of each of the at least one first pixel region, an initial set of confidence boxes is constructed, and the at least one first pixel region is stitched together to form an initial stitched pixel region.

[0323] Based on the coverage relationship between the second candidate regions of each of the at least one target pixel region and the set of confidence boxes, the set of confidence boxes and the stitched pixel regions are iteratively updated to obtain the segmentation result of the reference image; wherein, each iterative update process includes the following operations:

[0324] If a second candidate region covers multiple candidate regions in the previously updated confidence box set, then the second candidate region is added to the previously updated confidence box set, and the multiple candidate regions are removed from the previously updated confidence box set to obtain the current updated confidence box set.

[0325] The target pixel region corresponding to the second candidate region is stitched to the previously updated stitched pixel region to obtain the currently updated stitched pixel region.

[0326] Optionally, the identification module is further configured to:

[0327] If a second candidate region does not contain the plurality of candidate regions, then the target pixel region is determined as the newly added region compared to the last updated spliced ​​pixel region;

[0328] When the proportion of the newly added region in the target pixel region is greater than the second proportion threshold, the second candidate region is added to the previously updated confidence box set to obtain the current updated confidence box set; and the target pixel region is stitched to the previously updated stitched pixel region to obtain the current updated stitched pixel region.

[0329] In this embodiment, entity categories associated with a reference image are used as prompts to perform entity detection in the reference image, obtaining multiple entity candidate regions. The entity categories associated with the reference image are not limited to those labeled during the training phase, but may also include entity categories not labeled during the training phase. Therefore, when performing entity detection, entity candidate regions corresponding to entity categories labeled during the training phase and entity candidate regions corresponding to entity categories not labeled can be detected, thereby enabling the multiple entity candidate regions obtained to cover all entities to be identified as much as possible, improving the comprehensiveness of entity detection.

[0330] Secondly, from multiple entity candidate regions, a first candidate region with a confidence level greater than a first threshold (i.e., a high-confidence candidate region) is selected, while a second candidate region with a confidence level not greater than the first threshold (i.e., a low-confidence candidate region) is retained. Then, each first candidate region is segmented into pixels to obtain a corresponding first pixel region, thus identifying the pixel regions of salient entities in the reference image. Simultaneously, each second candidate region is segmented into pixels to obtain a corresponding second pixel region, and a target pixel region with a larger proportion in each second candidate region is selected to identify the pixel regions of missed entities. Finally, based on the obtained first pixel regions and target pixel regions, the segmentation result of the reference image is generated. This not only identifies salient entities in the reference image but also uses low-confidence candidate regions to supplement undetected entities, greatly reducing the number of missed entities and thus improving the accuracy of entity recognition in the image.

[0331] In addition, in the scene building task, low-confidence entity candidate regions are retained and used to supplement entities that were not successfully detected. In this way, even if there are differences between the image style and labeled entity categories of the dataset used during training and the actual image style and expected entity categories of the reference images in the scene building task, entity omission problems can be detected and corrected in time, thereby improving the performance of the scene building task.

[0332] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0333] Based on the same technical concept, this application provides a computer device, which may be the terminal device shown in FIG1 and / or a server. As shown in FIG11, it includes at least one processor 1101 and a memory 1102 connected to at least one processor. In this application embodiment, the specific connection medium between the processor 1101 and the memory 1102 is not limited. FIG11 shows the processor 1101 and the memory 1102 connected via a bus. The bus can be divided into address bus, data bus, control bus, etc.

[0334] In this embodiment of the application, the memory 1102 stores instructions that can be executed by at least one processor 1101. By executing the instructions stored in the memory 1102, at least one processor 1101 can perform the steps of the above-described image recognition method.

[0335] As one embodiment, the processor 1101 performs entity detection in the reference image based on the entity categories associated with the reference image, and obtains multiple entity candidate regions. Each entity candidate region corresponds to an entity category and the confidence level of that entity category. The confidence level represents the probability that an entity in the corresponding entity candidate region belongs to an entity category.

[0336] Processor 1101 selects at least one first candidate region and at least one second candidate region from multiple entity candidate regions, wherein the confidence level of the first candidate region is greater than a first threshold and the confidence level of the second candidate region is not greater than the first threshold; processor 1101 segments a first pixel region containing an entity from each first candidate region; and segments a second pixel region containing an entity from each second candidate region; processor 1101 selects at least one target pixel region from the obtained at least one second pixel region; wherein the proportion of each target pixel region in the corresponding second candidate region is greater than a first proportion threshold; processor 1101 obtains a segmentation result of a reference image based on at least one first pixel region and at least one target pixel region, and the segmentation result of the reference image is stored in memory 1102.

[0337] The processor 1101 is the control center of the computer device, capable of connecting to various parts of the computer device via various interfaces and lines. It achieves entity recognition by running or executing instructions stored in the memory 1102 and accessing data stored in the memory 1102. Optionally, the processor 1101 may include one or more processing units. The processor 1101 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and applications, while the modem processor primarily handles wireless communication. It is understood that the modem processor may not be integrated into the processor 1101. In some embodiments, the processor 1101 and the memory 1102 may be implemented on the same chip; in other embodiments, they may be implemented on separate chips.

[0338] Processor 1101 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0339] Memory 1102, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 1102 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 1102 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer device, but is not limited thereto. In the embodiments of this application, memory 1102 may also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0340] Based on the same inventive concept, embodiments of this application provide a computer-readable storage medium storing a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform the steps of the above-described image recognition method.

[0341] Based on the same inventive concept, this application provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, which, when executed by a computer device, cause the computer device to perform the steps of the above-described image recognition method.

[0342] Those skilled in the art will understand that embodiments of the present invention can be provided as methods or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0343] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer apparatus or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0344] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer device or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0345] These computer program instructions may also be loaded onto a computer device or other programmable data processing device to cause a series of operational steps to be performed on the computer device or other programmable device to produce a process implemented by the computer device, such that the instructions, which execute on the computer device or other programmable device, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0346] In summary, this application provides an image recognition method, apparatus, device, computer-readable storage medium, and computer program product. By using entity categories associated with a reference image as cues, entity detection is performed in the reference image. Because the associated entity categories are not limited to those labeled during training, the detected candidate regions can cover as many expected entities as possible, improving the comprehensiveness of entity detection. From the multiple candidate regions, first and second candidate regions with different confidence levels are selected. First and second pixel regions are segmented respectively, and target pixel regions are selected from the second pixel regions. Finally, the segmentation result of the reference image is obtained based on these pixel regions. This method not only identifies salient entities in the reference image but also uses low-confidence second candidate regions to supplement missed entities, reducing the number of missed entities and improving the accuracy of image entity recognition. In scene construction tasks, it can effectively address the differences between the training dataset and the actual reference image, promptly correcting entity omissions and improving the accuracy and efficiency of scene construction tasks.

[0347] Furthermore, feature extraction is performed on the reference image to obtain image features. Based on these features, the scene category and multiple preliminary descriptive words of the reference image are obtained. These are then combined with supplementary descriptive words selected from a pre-defined word library that match the scene category to obtain multiple entity prompts. The use of an image description model to automatically generate preliminary descriptive words improves the efficiency of prompt acquisition. Hot word libraries for different scenes can supplement common and easily overlooked prompts, making entity prompts more comprehensive. This helps to more accurately determine the entity category associated with the reference image, thereby improving the accuracy and efficiency of entity detection and reducing detection errors.

[0348] Furthermore, the topic descriptors of the reference image are obtained and expanded to obtain multiple target descriptors. These target descriptors are then combined with the initial and supplementary descriptors to obtain entity prompts. By expanding the topic descriptors, the diversity and flexibility of the prompt source are increased, enabling the entity prompts to more comprehensively reflect the entity information in the reference image and improving the accuracy of entity category recognition during entity detection.

[0349] Furthermore, the target descriptor, preliminary descriptor, and supplementary descriptor are constructed into a descriptor set, which undergoes preprocessing operations including deduplication of prompt words, removal of erroneous words, supplementation with synonyms, replacement of ambiguous words, and supplementation with invalid words. Deduplication of prompt words reduces the amount of data for subsequent processing and lowers computational complexity; removal of erroneous words avoids generating inaccurate or broad detection boxes; supplementation with synonyms improves the accuracy of descriptions and reduces the possibility of missed detections; replacement of ambiguous words reduces post-processing errors caused by semantic ambiguity; and supplementation with invalid words avoids outputting invalid entities as valid word categories, improving the accuracy and reliability of entity detection.

[0350] Furthermore, reference prompts that meet the preset segmentation granularity are selected from multiple entity prompts for entity detection to obtain reference candidate regions. These are then combined with other entity prompts to obtain supplementary candidate regions, and duplicate supplementary candidate regions are removed. This approach prioritizes the identification of entities that are easily fragmented or ambiguous, avoiding duplicate entity detection and the generation of low-quality candidate regions. This improves the accuracy and efficiency of entity detection, making the detection results more consistent with actual needs.

[0351] Furthermore, features are extracted and fused from the reference prompts and reference images respectively to obtain target fused features and preliminary candidate regions. Deduplication is then performed to obtain reference candidate regions. By combining features from text and image modalities, the semantic information of different modalities complements and collaborates with each other, compensating for the deficiencies of single-modal features, improving the accuracy and reliability of reference candidate regions, and enabling the detection results to more accurately reflect the entity information in the image.

[0352] Furthermore, the preliminary candidate regions are clustered according to the similarity of their corresponding entity categories. Each cluster is then deduplicated to obtain reference candidate regions. Clustering deduplication effectively detects instances where similar categories correspond to the same entity, improving deduplication accuracy and avoiding duplicate detections. Simultaneously, the enclosing relationships between candidate regions are considered to prevent duplication, partial omissions, or fragmentation of entity detection results, making the detection results more complete and accurate.

[0353] Furthermore, when selecting the second pixel region, the target pixel region is determined based on its proportion within the corresponding second candidate region, the size of the second candidate region, and its confidence level. Quality checks are performed on the second candidate region and the second pixel region, filtering out candidate regions and pixel regions with low quality or containing incomplete entities. This improves the effectiveness of subsequent supplementation of missed entities, making the final segmentation result more accurate and reliable.

[0354] Furthermore, an initial set of confidence boxes and a stitched pixel region are constructed based on the first pixel region, and iteratively updated according to the coverage relationship between the second candidate region and the confidence box set. If the second candidate region covers multiple candidate regions, it replaces the original multiple candidate regions, avoiding fragmentation of entities and improving the accuracy of object detection and segmentation. At the same time, low-confidence candidate regions are newly validated to supplement missed entity candidate regions, reducing the number of missed entities and making the segmentation result more complete.

[0355] Furthermore, if the second candidate region does not contain multiple candidate regions, the decision to add the second candidate region to the confidence box set is determined based on the proportion of the newly added region compared to the stitched pixel region. Identifying missed entities based on the area of ​​the newly added region and supplementing them to the segmentation result reduces the number of missed entities, improves the accuracy of image entity recognition, and makes the segmentation result more reflective of the real entity information in the image.

[0356] Furthermore, after obtaining the segmentation results of the reference image, an augmentation check is performed on at least one first pixel region and at least one target pixel region. The confidence box set and the stitched pixel region are updated by determining the coverage relationship between the second candidate region and the candidate regions in the confidence box set. If the second candidate region covers multiple candidate regions, it indicates that it represents a complete entity, and the second candidate region replaces the original multiple candidate regions, while simultaneously replacing the corresponding pixel regions in the stitched pixel region. This process further ensures the integrity of the entity, avoids over-segmentation, and improves the accuracy of image recognition. Simultaneously, by checking the connectivity of pixel regions, the integrity of the entity is further verified, making the recognition results more reliable.

[0357] In scene construction, scene entities are separated from the segmentation results of the reference image, and the remaining entities are divided into main bodies and decorative bodies based on the pixel area. This classification method allows different types of entities to be processed and utilized more rationally, improving the accuracy and efficiency of subsequent scene construction. Different types of entities can be connected to suitable downstream systems for processing, such as main bodies to an entity recommendation system, decorative bodies to a decoration and beautification system, and scene entities to a terrain generation system, making scene construction more professional and efficient.

[0358] In scene building tasks, maximum and minimum entity counts are set for reference images. Setting the maximum entity count filters out excessive small entities that have little impact on scene building, reducing unnecessary calculations and processing, improving the processing efficiency of downstream tasks, and making the scene building results more concise and targeted. Setting the minimum entity count ensures scene richness, avoids overly monotonous scenes, and improves the quality of scene building.

[0359] In scene building tasks, entities are sorted according to pixel area or placement relationship. Sorting by pixel area prioritizes important entities, improving scene building efficiency. Sorting by placement relationship ensures that entity placement conforms to actual logic, improving the rationality and accuracy of scene building. Furthermore, if there are entity prompts input from upstream, their corresponding entities are prioritized for display, further enhancing scene building accuracy and user satisfaction.

[0360] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0361] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An image recognition method, executed by a computer device, comprising: Based on the entity categories associated with the reference image, entity detection is performed in the reference image to obtain multiple entity candidate regions. Each entity candidate region corresponds to an entity category and the confidence level of that entity category. At least one first candidate region and at least one second candidate region are selected from the plurality of entity candidate regions, wherein the confidence level of the first candidate region is greater than a first threshold and the confidence level of the second candidate region is not greater than the first threshold; Segment the first pixel region containing the entity from each first candidate region; And, from each second candidate region, segment out the second pixel region containing the entity; At least one target pixel region is selected from at least one second pixel region obtained; wherein, the proportion of each target pixel region in the second candidate region from which the target pixel region is segmented is greater than a first proportion threshold; and The segmentation result of the reference image is obtained based on the at least one first pixel region and the at least one target pixel region.

2. The method of claim 1, wherein the entity categories are obtained in the following manner: Feature extraction is performed on the reference image to obtain the corresponding image features; Based on the image features, the scene category and multiple preliminary descriptive words of the reference image are obtained; each preliminary descriptive word is used to describe an entity initially identified from the reference image; From multiple preset word libraries, a target word library matching the scene category is selected, wherein... The target lexicon includes multiple supplementary descriptive terms; each supplementary descriptive term is used to describe an entity that occurs more frequently than a preset threshold under the scene category. Based on the multiple preliminary descriptive words and the multiple supplementary descriptive words, multiple entity prompt words are obtained; wherein, each entity prompt word represents an entity category.

3. The method as described in claim 2, wherein obtaining multiple entity prompt words based on the plurality of preliminary descriptive words and the plurality of supplementary descriptive words includes: Obtain the topic descriptive words corresponding to the reference image settings; Based on the semantic information contained in the topic descriptors, the topic descriptors are expanded to obtain multiple target descriptors; Based on the multiple target descriptors, the multiple preliminary descriptors, and the multiple supplementary descriptors, multiple entity prompt words are obtained.

4. The method as described in claim 3, wherein obtaining multiple entity prompt words based on the plurality of target descriptors, the plurality of preliminary descriptors, and the plurality of supplementary descriptors includes: Based on the multiple target descriptors, the multiple preliminary descriptors, and the multiple supplementary descriptors, a descriptor set is constructed; Perform a preprocessing operation on the descriptor set, and use multiple descriptors in the preprocessed descriptor set as the multiple entity prompt words; The preprocessing operation includes at least one of the following: deduplication of prompt words; Error word removal; synonym replacement; ambiguous word replacement; invalid word replacement.

5. The method according to any one of claims 2 to 4, wherein the entity detection in the reference image based on the entity categories associated with the reference image to obtain multiple entity candidate regions includes: Select multiple reference prompt words that meet the preset segmentation granularity from the multiple entity prompt words; Based on the multiple reference prompt words, entity detection is performed in the reference image to obtain multiple reference candidate regions; Furthermore, based on other entity prompts among the plurality of entity prompts, entity detection is performed in the reference image to obtain a plurality of supplementary candidate regions; For each supplementary candidate region, the following steps are performed: if the intersection-union ratio between a supplementary candidate region and any of the reference candidate regions is greater than the first decision threshold, then the supplementary candidate region is removed. At least one supplementary candidate region and the plurality of reference candidate regions are retained as the plurality of entity candidate regions.

6. The method of claim 5, wherein the step of performing entity detection in the reference image based on the plurality of reference cue words to obtain a plurality of reference candidate regions includes: Feature extraction is performed on the multiple reference prompt words to obtain the corresponding text features; In addition, feature extraction is performed on the reference image to obtain the corresponding image features; The text features and the image features are fused to obtain the target fused features; Based on the target fusion features, multiple preliminary candidate regions in the reference image are obtained; The multiple preliminary candidate regions are deduplicated to obtain the multiple reference candidate regions.

7. The method of claim 6, wherein deduplicating the plurality of preliminary candidate regions to obtain the plurality of reference candidate regions comprises: Based on the similarity between the entity categories corresponding to each of the multiple preliminary candidate regions, the multiple preliminary candidate regions are clustered to obtain multiple candidate region sets; For each candidate region set, perform the following: Select the baseline candidate region with the highest confidence from the candidate region set; Remove each other preliminary candidate region in the candidate region set whose intersection-union ratio with the benchmark candidate region is greater than the second decision threshold, or remove each other preliminary candidate region in the candidate region set that is located inside the benchmark candidate region; Based on at least one preliminary candidate region retained by each of the multiple candidate region sets, the multiple reference candidate regions are obtained.

8. The method of any one of claims 1 to 7, wherein filtering at least one target pixel region from the obtained at least one second pixel region comprises: For each second pixel region, perform the following: If the proportion of a second pixel region in the corresponding second candidate region is greater than the first proportion threshold, and the size of the second candidate region is within a preset range, and the confidence level of the second candidate region is greater than the second threshold, then the second pixel region is determined to be the target pixel region.

9. The method according to any one of claims 1 to 7, wherein obtaining the segmentation result of the reference image based on the at least one first pixel region and the at least one target pixel region includes: Based on the first candidate regions of each of the at least one first pixel region, an initial set of confidence boxes is constructed; And, the at least one first pixel region is stitched together to form an initial stitched pixel region; Based on the coverage relationship between the second candidate regions of each of the at least one target pixel region and the set of confidence boxes, the set of confidence boxes and the stitched pixel regions are iteratively updated to obtain the segmentation result of the reference image; wherein, each iterative update process includes the following operations: If a second candidate region covers multiple candidate regions in the previously updated confidence box set, then the second candidate region is added to the previously updated confidence box set, and the multiple candidate regions are removed from the previously updated confidence box set to obtain the current updated confidence box set. The target pixel region corresponding to the second candidate region is stitched to the previously updated stitched pixel region to obtain the currently updated stitched pixel region.

10. The method of claim 9, further comprising: If a second candidate region does not contain the plurality of candidate regions, then the target pixel region is determined as the newly added region compared to the last updated spliced ​​pixel region; When the proportion of the newly added region in the target pixel region is greater than the second proportion threshold, the second candidate region is added to the previously updated confidence box set to obtain the current updated confidence box set. And, the target pixel region is stitched to the previously updated stitched pixel region to obtain the currently updated stitched pixel region.

11. An image recognition device, comprising: The detection module is used to perform entity detection in the reference image based on the entity categories associated with the reference image, and obtain multiple entity candidate regions, each entity candidate region corresponding to an entity category and the confidence level of that entity category; The filtering module is used to filter at least one first candidate region and at least one second candidate region from the plurality of entity candidate regions, wherein the confidence level of the first candidate region is greater than a first threshold and the confidence level of the second candidate region is not greater than the first threshold. The segmentation module is used to segment the first pixel region containing the entity from each first candidate region; And, from each second candidate region, segment out the second pixel region containing the entity; The filtering module is further configured to filter out at least one target pixel region from the obtained at least one second pixel region; wherein the proportion of each target pixel region in the corresponding second candidate region is greater than a first proportion threshold. and The recognition module is used to obtain the segmentation result of the reference image based on the at least one first pixel region and the at least one target pixel region.

12. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any one of claims 1 to 10.

13. A computer-readable storage medium storing a computer program executable by a computer device, which, when run on the computer device, causes the computer device to perform the steps of the method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program stored on a computer-readable storage medium, the computer program including program instructions that, when executed by a computer device, cause the computer device to perform the steps of the method according to any one of claims 1-10.