Entity recognition method and device, electronic equipment, storage medium and program product
By segmenting the image and identifying vector images to be recognized, and entity recognition combined with entity regions and element categories, the problem of low entity recognition accuracy in the prior art is solved, and higher recognition accuracy is achieved.
Patent Information
- Application Number
- CN202311584762.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-23
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, the accuracy of entity recognition is low because the image features to be identified can only be described in the overall macroscopic image.
By obtaining the vector image to be identified, image segmentation is performed to obtain the entity area, and element recognition is performed on vector elements, and entity recognition is performed on combining entity areas and element categories.
The accuracy of entity recognition is improved, making the resulting vector recognition results more accurate.
Smart Images

Figure CN120032153A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an entity recognition method, device, electronic device, storage medium and program product. Background Art
[0002] Artificial Intelligence (AI) is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model is also called a large model or a basic model. After fine-tuning, it can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0003] In the related art, for entity recognition, usually features are directly extracted from the image to be recognized. After the features of the image to be recognized are obtained, the entity recognition model is called, and entity recognition is performed on the image to be recognized based on the features of the image to be recognized to obtain an entity recognition result. Since the features of the image to be recognized can only describe the image to be recognized from an overall macro perspective, the accuracy of the obtained entity recognition result is relatively low. Summary of the invention
[0004] The embodiments of the present application provide an entity recognition method, device, electronic device, computer-readable storage medium, and computer program product, which can effectively improve the accuracy of entity recognition.
[0005] The technical solution of the embodiment of the present application is implemented as follows:
[0006] The present application provides an entity recognition method, including:
[0007] Acquire a vector graph to be identified including at least one vector entity, wherein the vector entity includes at least one vector element;
[0008] Performing image segmentation on the vector image to be identified to obtain at least one entity region in the vector image to be identified, wherein the entity region includes at least one of the vector entities;
[0009] Performing element recognition on each of the vector elements in the vector image to be recognized to obtain entity categories corresponding to each of the vector elements;
[0010] In combination with the entity area and the entity categories corresponding to each of the vector elements, entity recognition is performed on the vector image to be identified to obtain an entity recognition result, and the entity recognition result is used to indicate the entity categories corresponding to each of the vector entities in the vector image to be identified.
[0011] The present application provides an entity recognition device, including:
[0012] An acquisition module, used for acquiring a vector graph to be identified including at least one vector entity, wherein the vector entity includes at least one vector element;
[0013] An image segmentation module, configured to perform image segmentation on the vector image to be identified, and obtain at least one entity region in the vector image to be identified, wherein the entity region includes at least one of the vector entities;
[0014] An element recognition module is used to perform element recognition on each of the vector elements in the vector image to be recognized, and obtain entity categories corresponding to each of the vector elements;
[0015] The entity recognition module is used to perform entity recognition on the vector map to be recognized in combination with the entity area and the entity categories corresponding to each of the vector elements to obtain an entity recognition result, wherein the entity recognition result is used to indicate the entity categories corresponding to each of the vector entities in the vector map to be recognized.
[0016] In the above scheme, the above-mentioned image segmentation module is also used to perform pixel point identification on each pixel point in the vector map to be identified, and obtain pixel identification results corresponding to each pixel point; when the pixel identification result indicates that the pixel point is a pixel point within the area where the vector element in the vector map to be identified is located, the pixel point is determined as a target pixel point; clustering processing is performed on the target pixel points in the vector map to be identified to obtain at least one initial entity area in the vector map to be identified, and the target pixel points in the initial entity area belong to the same vector element; based on the initial entity area, at least one entity area in the vector map to be identified is determined.
[0017] In the above scheme, the above-mentioned image segmentation module is also used to, when the number of the initial entity regions is one, determine the initial entity region as the entity region; when the number of the initial entity regions is multiple, cluster the initial entity regions in the vector map to be identified to obtain at least one reference entity region in the vector map to be identified, and the initial entity regions in the reference entity regions belong to the same vector entity; based on the reference entity region, determine at least one of the entity regions in the vector map to be identified.
[0018] In the above scheme, the above-mentioned image segmentation module is also used to, when the number of the reference entity areas is one, determine the reference entity area as the entity area; when the number of the reference entity areas is multiple, obtain the value of the regional morphological parameters of each of the reference entity areas, and the value range of the morphological parameters of the vector entity; compare the value of each of the regional morphological parameters with the value range of the morphological parameter, respectively, to obtain the parameter comparison result of each of the regional morphological parameters; when the reference comparison result indicates that the value of the regional morphological parameter is within the value range of the morphological parameter, determine the reference entity area corresponding to the regional morphological parameter as the entity area.
[0019] In the above scheme, the above-mentioned element recognition is realized through an element recognition network, and the element recognition network includes a recognition layer, a feature extraction layer and a classification layer. The above-mentioned element recognition module is also used to call the recognition layer to perform vector recognition on the vector image to be recognized, and obtain each of the vector elements in the vector image to be recognized; call the feature extraction layer to perform feature extraction on each of the vector elements respectively, and obtain the element features corresponding to each of the vector elements; for each of the element features, call the classification layer to classify the corresponding vector elements based on the element features, and obtain the entity category corresponding to the vector element.
[0020] In the above scheme, the above-mentioned entity recognition module is also used to obtain the regional confidence of each of the entity regions and the element confidence of each of the vector elements, the element confidence is used to indicate the probability that the vector element is the corresponding entity category, and the regional confidence is used to indicate the probability that the number of the vector entities in the entity region is one; summing up the regional confidences to obtain a first confidence, and summing up the element confidences to obtain a second confidence; combining the first confidence and the second confidence, performing entity recognition on the vector image to be identified to obtain the entity recognition result.
[0021] In the above scheme, the above entity recognition module is also used to compare the first confidence and the first confidence threshold to obtain a first comparison result, and compare the second confidence and the second confidence threshold to obtain a second comparison result; when the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, based on the entity category corresponding to the vector element and the entity area, the vector image to be recognized is subjected to entity recognition to obtain the entity recognition result; when the first comparison result indicates that the first confidence is less than the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, based on the entity category corresponding to each of the vector elements, the vector image to be recognized is subjected to entity recognition to obtain the entity recognition result; when the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is less than the second confidence threshold, based on the entity area, the vector image to be recognized is subjected to entity recognition to obtain the entity recognition result.
[0022] In the above scheme, the entity recognition result includes the entity recognition results corresponding to each of the entity areas, and the above entity recognition module is also used to perform the following processing for each of the entity areas: determining the vector elements located in the entity area as the target vector elements corresponding to the entity area; when the number of the target vector elements is one, determining the entity category corresponding to the target vector element as the entity recognition result corresponding to the entity area; when the number of the target vector elements is multiple, clustering the target vector elements based on the entity categories corresponding to the target vector elements, to obtain at least one target vector entity in the entity area, and the entity categories corresponding to each of the target vector elements in the target vector entity are the same; determining the entity categories of the target vector elements in each of the target vector entities as the entity categories of the corresponding target vector entities; determining the entity category of each of the target vector entities as the entity recognition result corresponding to the entity area.
[0023] In the above scheme, the above-mentioned entity recognition module is also used to, when the number of the vector elements in the vector image to be recognized is one, determine the entity category corresponding to the vector element as the entity recognition result; when the number of the vector elements in the vector image to be recognized is multiple, based on the entity categories respectively corresponding to the vector elements, cluster the vector elements in the vector image to be recognized to obtain at least one vector entity in the vector image to be recognized, and the entity categories respectively corresponding to the vector elements in the vector entity are the same; determine the entity categories of the vector elements in each of the vector entities as the entity categories of the corresponding vector entities; and determine the entity categories of each of the vector entities as the entity recognition result.
[0024] In the above scheme, the entity recognition result includes the entity recognition results corresponding to each of the entity areas, and the above entity recognition module is also used to perform the following processing for each of the entity areas: determining the vector elements located in the entity area as the target vector elements corresponding to the entity area; when the number of the target vector elements is one, determining the entity category corresponding to the target vector element as the entity recognition result corresponding to the entity area; when the number of the target vector elements is multiple, obtaining the expected entity contour, and based on the expected entity contour, clustering the target vector elements to obtain at least one target vector entity in the entity area, and the contour formed by each of the target vector elements in the target vector entity satisfies the expected entity contour; determining the entity category corresponding to the expected entity contour as the entity recognition result corresponding to the entity area.
[0025] In the above scheme, the above entity recognition module is also used to perform the following processing for each of the entity areas: determine the vector elements located in the entity area as the target vector elements corresponding to the entity area; when the number of the target vector elements is one, determine the entity category corresponding to the target vector element as the entity recognition result corresponding to the entity area; when the number of the target vector elements is multiple, cluster the target vector elements based on the entity categories respectively corresponding to the target vector elements to obtain at least one target vector entity in the entity area, and the entity categories respectively corresponding to the target vector elements in the target vector entity are the same; determine the entity categories of the target vector elements in each of the target vector entities as the entity categories of the corresponding target vector entities; determine the entity categories of each of the target vector entities as the entity recognition result corresponding to the entity area.
[0026] An embodiment of the present application provides an electronic device, including:
[0027] A memory for storing computer executable instructions or computer programs;
[0028] The processor is used to implement the entity recognition method provided in the embodiment of the present application when executing the computer executable instructions or computer program stored in the memory.
[0029] An embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions for causing a processor to execute and implement the entity recognition method provided in the embodiment of the present application.
[0030] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, so that the electronic device performs the entity recognition method described in the embodiment of the present application.
[0031] The embodiments of the present application have the following beneficial effects:
[0032] By obtaining a vector image to be identified including at least one vector entity, performing image segmentation on the vector image to be identified, obtaining at least one entity region in the vector image to be identified, and performing element recognition on each vector element in the vector image to be identified, obtaining the entity category corresponding to each vector element, and combining the entity region and the entity category corresponding to the vector element, performing entity recognition on the vector image to be identified, and obtaining an entity recognition result. In this way, by performing image segmentation on the vector image to be identified, obtaining at least one entity region in the vector image to be identified, the region including the vector entity in the vector image to be identified is identified from the overall macroscopic perspective, and by performing element recognition on each vector element in the vector image to be identified, the entity category corresponding to each vector element is obtained, thereby performing element recognition on the vector elements in the vector image to be identified from the dimensions of each vector element constituting the vector entity and from the detailed microscopic perspective in advance, and performing entity recognition on the vector image to be identified from the overall macroscopic entity region and the detailed microscopic entity category corresponding to the vector element, so that the obtained vector recognition result is more accurate, and the accuracy of entity recognition is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] Figure 1 It is a schematic diagram of the architecture of the entity recognition system provided in the embodiment of the present application;
[0034] Figure 2 is a schematic diagram of the structure of an electronic device for entity identification provided in an embodiment of the present application;
[0035] Figure 3This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 1 ;
[0036] Figure 4 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 2 ;
[0037] Figure 5 is a schematic diagram of the structure of a pixel recognition network provided in an embodiment of the present application;
[0038] Figure 6 It is a schematic diagram of the principle of the element recognition network provided in the embodiment of the present application;
[0039] Figure 7 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 3 ;
[0040] Figure 8 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 4 ;
[0041] Fig. 9 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 5 ;
[0042] Fig.10 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 6 ;
[0043] Fig.11 This is a schematic diagram of the effect of the vector diagram to be identified provided in the embodiment of the present application;
[0044] Fig.12 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 7 ;
[0045] Fig.13 It is a schematic diagram of the principle of the entity area provided in the embodiment of the present application. DETAILED DESCRIPTION
[0046] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. The described embodiments should not be regarded as limiting the present application. All other embodiments obtained by ordinary technicians in the field without making creative work are within the scope of protection of this application.
[0047] In the following description, reference is made to “some embodiments”, which describe a subset of all possible embodiments, but it will be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0048] In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged with a specific order or sequence where permitted, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.
[0049] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0050] Before further describing the embodiments of the present application in detail, the nouns and terms involved in the embodiments of the present application are explained. The nouns and terms involved in the embodiments of the present application are subject to the following interpretations.
[0051] 1) Artificial Intelligence (AI): It is an interdisciplinary subject that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model is also called a large model or a basic model. After fine-tuning, it can be widely used in downstream tasks in various major directions of artificial intelligence. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0052] 2) Computer Vision Technology (CV): Computer vision is a science that studies how to make machines "see". To put it more specifically, it refers to machine vision such as using cameras and computers to replace human eyes to identify and measure targets, and further perform image processing to make computer processing into images that are more suitable for human eye observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Large model technology has brought important changes to the development of computer vision technology. Pre-trained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and fingerprint recognition.
[0053] 3) Machine Learning (ML): It is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning by induction. The pre-trained model is the latest development in deep learning, which integrates the above technologies.
[0054] 4) Convolutional Neural Networks (CNN): It is a type of feed forward neural network (FNN) that includes convolution calculations and has a deep structure. It is one of the representative algorithms of deep learning. Convolutional neural networks have representation learning capabilities and can perform shift-invariant classification on input images according to their hierarchical structure.
[0055] 5) In response to: used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed may be in real time or have a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations executed are executed.
[0056] 6) Vector Graphics: Graphics described by straight lines and curves. The elements that make up these graphics are points, lines, rectangles, polygons, circles, arcs, etc. They are all calculated through mathematical formulas and have the characteristic of not being distorted after editing. For example, the vector graphics of a painting are actually the outline of the frame formed by line segments, and the color of the frame and the color enclosed by the frame determine the color of the painting. Vector graphics are also called object-oriented images or drawing images. In the traditional Chinese version, they are called vector graphics. In computer graphics, geometric primitives such as points, straight lines or polygons based on mathematical equations are used to represent images. The biggest advantage of vector graphics is that they will not be distorted regardless of enlargement, reduction or rotation; the biggest disadvantage is that it is difficult to express realistic image effects with rich color levels.
[0057] 7) Image segmentation: In the field of computer vision, image segmentation refers to the process of subdividing a digital image into multiple image sub-regions (collections of pixels) (also called superpixels). The purpose of image segmentation is to simplify or change the representation of an image so that it is easier to understand and analyze. Image segmentation is often used to locate objects and boundaries (lines, curves, etc.) in an image.
[0058] 8) Convolutional Neural Networks (CNN): It is a type of feed forward neural network (FNN) that includes convolution calculations and has a deep structure. It is one of the representative algorithms of deep learning. Convolutional neural networks have representation learning capabilities and can perform shift-invariant classification on input images according to their hierarchical structure.
[0059] 9) Convolutional layer: Each convolutional layer in a convolutional neural network consists of several convolutional units, and the parameters of each convolutional unit are optimized through the back propagation algorithm. The purpose of the convolution operation is to extract different features of the input. The first convolutional layer may only extract some low-level features such as edges, lines, and corners. More layers of the network can iteratively extract more complex features from low-level features.
[0060] 10) Pooling layer: After the convolution layer performs feature extraction, the output feature map is passed to the pooling layer for feature selection and information filtering. The pooling layer contains a pre-set pooling function, which replaces the result of a single point in the feature map with the feature map statistics of its adjacent area. The pooling layer selects the pooling area in the same way as the convolution kernel scans the feature map, which is controlled by the pooling size, step size, and padding.
[0061] 11) Fully-Connected Layer: The fully connected layer in a convolutional neural network is equivalent to the hidden layer in a traditional feedforward neural network. The fully connected layer is located at the end of the hidden layer of the convolutional neural network and only transmits signals to other fully connected layers. The feature map loses its spatial topological structure in the fully connected layer, is expanded into a vector and passes through the activation function.
[0062] During the implementation of the embodiments of the present application, the applicant discovered that the related technology has the following problems:
[0063] In the related art, for entity recognition, usually features are directly extracted from the image to be recognized. After the features of the image to be recognized are obtained, the entity recognition model is called, and entity recognition is performed on the image to be recognized based on the features of the image to be recognized to obtain an entity recognition result. Since the features of the image to be recognized can only describe the image to be recognized from an overall macro perspective, the accuracy of the obtained entity recognition result is relatively low.
[0064] The embodiments of the present application provide an entity recognition method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can effectively improve the accuracy of entity recognition. The following describes an exemplary application of the entity recognition system provided by the embodiments of the present application.
[0065] See also Figure 1 , Figure 1 It is a schematic diagram of the architecture of the entity recognition system 100 provided in an embodiment of the present application. The terminal (terminal 400 is shown as an example) is connected to the server 200 via the network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0066] The terminal 400 is used for the user to use the client 410, and the entity recognition result is displayed on the graphical interface 410-1 (graphical interface 410-1 is shown as an example). The terminal 400 and the server 200 are connected to each other via a wired or wireless network.
[0067] In some embodiments, the server 200 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms. The terminal 400 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart TV, a smart watch, a car terminal, etc., but is not limited thereto. The electronic device provided in the embodiment of the present application may be implemented as a terminal or as a server. The terminal and the server may be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiment of the present application.
[0068] In some embodiments, the server 200 performs image segmentation on the vector image to be identified to obtain an entity area, and performs element recognition on each vector element in the vector image to be identified to obtain an entity category corresponding to each vector element, and combines the entity area and the entity category to perform entity recognition on the vector image to be identified to obtain an entity recognition result, and sends the entity recognition result to the terminal 400.
[0069] In other embodiments, the terminal 400 performs image segmentation on the vector image to be identified to obtain an entity region, and performs element recognition on each vector element in the vector image to be identified to obtain an entity category corresponding to each vector element, and performs entity recognition on the vector image to be identified in combination with the entity region and the entity category to obtain an entity recognition result, and sends the entity recognition result to the server 200.
[0070] In other embodiments, the embodiments of the present application can be implemented with the aid of cloud technology. Cloud technology refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing.
[0071] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology based on the cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud computing technology will become an important support. The background services of the technical network system require a large amount of computing and storage resources.
[0072] See also Figure 2 , Figure 2 is a schematic diagram of the structure of an electronic device 500 for entity identification provided in an embodiment of the present application, wherein: Figure 2 The electronic device 500 shown may be Figure 1The server 200 or the terminal 400 in Figure 2 The electronic device 500 shown includes: at least one processor 430, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together via a bus system 440. It is understood that the bus system 440 is used to achieve connection and communication between these components. In addition to the data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, the bus system 440 is not described in detail. Figure 2 Various buses are labeled as bus system 440 .
[0073] The processor 430 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., where the general-purpose processor can be a microprocessor or any conventional processor, etc.
[0074] The memory 450 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, etc. The memory 450 may optionally include one or more storage devices that are physically remote from the processor 430.
[0075] The memory 450 includes a volatile memory or a non-volatile memory, and may also include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), and the volatile memory may be a random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0076] In some embodiments, memory 450 can store data to support various operations, examples of which include programs, modules, and data structures, or a subset or superset thereof, as exemplarily described below.
[0077] Operating system 451, including system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, etc., for implementing various basic services and processing hardware-based tasks;
[0078] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420. Exemplary network interfaces 420 include: Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB).
[0079] In some embodiments, the entity identification device provided in the embodiments of the present application can be implemented in a software manner. Figure 2 The entity recognition device 455 stored in the memory 450 is shown, which can be software in the form of a program and a plug-in, etc., including the following software modules: an acquisition module 4551, an image segmentation module 4552, an element recognition module 4553, and an entity recognition module 4554. These modules are logical, so they can be arbitrarily combined or further split according to the functions implemented. The functions of each module will be described below.
[0080] In other embodiments, the entity identification device provided in the embodiments of the present application can be implemented in hardware. As an example, the entity identification device provided in the embodiments of the present application can be a processor in the form of a hardware decoding processor, which is programmed to execute the entity identification method provided in the embodiments of the present application. For example, the processor in the form of a hardware decoding processor can adopt one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field programmable gate arrays (FPGAs), or other electronic components.
[0081] In some embodiments, the terminal or server can implement the entity recognition method provided by the embodiment of the present application by running a computer program or a computer executable instruction. For example, the computer program can be a native program (e.g., a dedicated entity recognition program) or a software module in the operating system, for example, an entity recognition module that can be embedded in any program (such as an instant messaging client, an album program, an electronic map client, a navigation client); for example, it can be a local (Native) application (APP, Application), that is, a program that needs to be installed in the operating system to run. In short, the above-mentioned computer program can be an application, module or plug-in in any form.
[0082] The entity identification method provided in the embodiment of the present application will be described in conjunction with the exemplary application and implementation of the server or terminal provided in the embodiment of the present application.
[0083] See also Figure 3 , Figure 3 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 1 , will combine Figure 3 Steps 101 to 104 are shown for illustration. The entity recognition method provided in the embodiment of the present application can be implemented by a server or a terminal alone, or by a server and a terminal in collaboration. The following description will be given by taking the server alone as an example.
[0084] In step 101, a vector image to be identified including at least one vector entity is obtained.
[0085] In some embodiments, a vector entity includes at least one vector element. Vector graphics are graphics (vector entities) described using straight lines and curves. The elements that constitute these vector entities are vector elements such as points, lines, rectangles, polygons, circles and arcs. They are all obtained by mathematical formula calculations and have the characteristic of not being distorted after editing. For example, the vector graphics of a painting are actually formed by line segments to form the outline of the outer frame, and the color of the outer frame and the color enclosed by the outer frame determine the color of the painting. Vector graphics are also called object-oriented images or drawing images. In the traditional Chinese version, they are called vector graphics. They are geometric graphics based on mathematical equations such as points, lines or polygons in computer graphics to represent images. The biggest advantage of vector graphics is that they will not be distorted regardless of enlargement, reduction or rotation; the biggest disadvantage is that it is difficult to express realistic image effects with rich color levels.
[0086] In some embodiments, the vector image to be identified may be a computer-aided design image (CAD-Computer Aided Design), also known as a CAD image, also known as a CAD construction drawing or a building plan, which is a drawing that uses AutoCAD software to create the overall layout of a project, the exterior shape, interior layout, structural structure, interior and exterior decoration, material methods, equipment, and construction of a building. CAD construction drawings have the characteristics of complete drawings, accurate expression, and specific requirements. They are the basis for engineering construction, preparation of construction drawing budgets, and construction organization design, and are also important technical documents for technical management.
[0087] As an example, the vector entities in the building plan can be buildings, electronic components, etc. When the vector entity is a building, the corresponding vector elements can be building walls, building bricks and tiles, etc. When the vector entity is an electronic component, the corresponding vector elements can be the constituent components of the electronic component.
[0088] In step 102, image segmentation is performed on the vector image to be identified to obtain at least one entity region in the vector image to be identified.
[0089] In some embodiments, the entity region includes at least one vector entity, image segmentation. In the field of computer vision, image segmentation refers to the process of subdividing a digital image into multiple image sub-regions (a collection of pixels) (also called superpixels). The purpose of image segmentation is to simplify or change the representation of an image so that the image is easier to understand and analyze. Image segmentation is often used to locate objects and boundaries (lines, curves, etc.) in an image.
[0090] In some embodiments, see Figure 4 , Figure 4 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 2 , Figure 3 The illustrated step 102 may be performed by Figure 4 Steps 1021 to 1024 are implemented as shown.
[0091] In step 1021, pixel recognition is performed on each pixel in the vector image to be recognized, and pixel recognition results corresponding to each pixel are obtained.
[0092] In some embodiments, the above pixel recognition result is used to indicate whether the pixel point is a pixel point within the region where the vector element in the vector image to be recognized is located.
[0093] In some embodiments, the above step 1021 can be implemented in the following manner: obtain a pixel recognition network, call the pixel recognition network, perform pixel recognition on each pixel in the vector graph to be recognized, and obtain pixel recognition results corresponding to each pixel.
[0094] In some embodiments, see Figure 5 , Figure 5 is a schematic diagram of the structure of a pixel recognition network provided in an embodiment of the present application, Figure 5 The pixel recognition network shown includes a feature encoding layer 51, a feature conversion layer 52 and a feature decoding layer 53, wherein the feature encoding layer 51, the feature conversion layer 52 and the feature decoding layer 53 include a convolutional layer, a pooling layer and a fully connected layer.
[0095] As an example, see Figure 5The above step 1021 can be implemented in the following manner: calling the feature encoding layer 51 to perform feature encoding on each pixel point in the vector diagram to be identified, and obtain the encoding results corresponding to each pixel point; calling the feature conversion layer 52 to perform feature conversion on each encoding result in the vector diagram to be identified, and obtain the feature conversion results corresponding to each pixel point; calling the feature decoding layer 53 to perform pixel point recognition on each pixel point in the vector diagram to be identified based on each feature conversion result, and obtain the pixel recognition results corresponding to each pixel point.
[0096] In step 1022, when the pixel recognition result indicates that the pixel point is a pixel point within the region where the vector element in the vector image to be recognized is located, the pixel point is determined as a target pixel point.
[0097] In some embodiments, the above step 1022 can be implemented as follows: for each pixel point, when the pixel recognition result corresponding to the pixel point indicates that the pixel point is a pixel point within the area where the vector element in the vector image to be identified is located, the pixel point is determined as a target pixel point.
[0098] In some embodiments, for each pixel point, when the pixel recognition result corresponding to the pixel point indicates that the pixel point is not a pixel point within the region where the vector element in the vector image to be recognized is located, the pixel point is not determined as a target pixel point.
[0099] In some embodiments, the target pixel points are pixel points within the region where the vector elements in the vector image to be identified are located.
[0100] In step 1023, clustering is performed on the target pixel points in the vector diagram to be identified to obtain at least one initial entity region in the vector diagram to be identified.
[0101] In some embodiments, the target pixels in the initial entity region belong to the same vector element, and the target pixels in different initial entity regions belong to different vector elements. The above clustering process refers to the process of clustering the initial entity regions belonging to the same vector element to obtain the corresponding initial entity regions.
[0102] As an example, the target pixel points in the vector diagram to be identified include target pixel point A, target pixel point B and target pixel point C, wherein target pixel point A belongs to vector element A1, target pixel point B belongs to vector element A2, and target pixel point C belongs to vector element A3. Then, the target pixel points in the vector diagram to be identified are clustered to obtain three initial entity areas in the vector diagram to be identified, wherein the initial entity area B1 includes the target pixel point A, the initial entity area B2 includes the target pixel point B, and the initial entity area B3 includes the target pixel point C.
[0103] As an example, the target pixel points in the vector diagram to be identified include target pixel point A, target pixel point B and target pixel point C, wherein target pixel point A belongs to vector element A1, target pixel point B belongs to vector element A1, and target pixel point C belongs to vector element A1. Then, the target pixel points in the vector diagram to be identified are clustered to obtain an initial entity area in the vector diagram to be identified, and the initial entity area B1 includes target pixel point A, target pixel point B and target pixel point C.
[0104] As an example, the target pixel points in the vector diagram to be identified include target pixel point A, target pixel point B and target pixel point C, wherein target pixel point A belongs to vector element A1, target pixel point B belongs to vector element A1, and target pixel point C belongs to vector element A2. Then, the target pixel points in the vector diagram to be identified are clustered to obtain two initial entity areas in the vector diagram to be identified, wherein the initial entity area B1 includes target pixel point A and target pixel point B, and the initial entity area B2 includes target pixel point C.
[0105] In step 1024, at least one entity region in the vector map to be identified is determined based on the initial entity region.
[0106] In some embodiments, the number of initial entity regions is less than or equal to the number of entity regions, and the above step 1024 can be implemented as follows: when the number of initial entity regions is one, the initial entity region is determined as the entity region; when the number of initial entity regions is multiple, the initial entity regions in the vector graph to be identified are clustered to obtain at least one reference entity region in the vector graph to be identified; based on the reference entity region, at least one entity region in the vector graph to be identified is determined.
[0107] In some embodiments, the initial entity regions in the reference entity regions belong to the same vector entity.
[0108] In some embodiments, the clustering process refers to a process of clustering initial entity regions belonging to the same vector entity to obtain corresponding reference entity regions.
[0109] Continuing with the example, the initial entity area includes initial entity area B1, initial entity area B2 and initial entity area B3, wherein initial entity area B1 belongs to vector entity W1, initial entity area B2 belongs to vector entity W1, and initial entity area B3 belongs to vector entity W2. Then, the initial entity areas in the vector diagram to be identified are clustered to obtain two reference entity areas in the vector diagram to be identified. The reference entity area C1 includes the initial entity area B1 and the initial entity area B2, and the reference entity area C2 includes the initial entity area B3.
[0110] Continuing with the example, the initial entity area includes the initial entity area B1, the initial entity area B2 and the initial entity area B3, wherein the initial entity area B1 belongs to the vector entity W1, the initial entity area B2 belongs to the vector entity W1, and the initial entity area B3 belongs to the vector entity W1. Then, the initial entity areas in the vector image to be identified are clustered to obtain a reference entity area in the vector image to be identified, and the reference entity area C1 includes the initial entity area B1, the initial entity area B2 and the initial entity area B3.
[0111] In some embodiments, the above-mentioned determination of at least one entity area in the vector image to be identified based on the reference entity area can be achieved in the following manner: when the number of reference entity areas is one, the reference entity area is determined as the entity area; when the number of reference entity areas is multiple, the values of the regional morphological parameters of each reference entity area and the value range of the morphological parameters of the vector entity are obtained; the values of the morphological parameters of each area are compared with the value range of the morphological parameters respectively to obtain the parameter comparison results of the morphological parameters of each area; when the reference comparison result indicates that the value of the regional morphological parameter is within the value range of the morphological parameter, the reference entity area corresponding to the regional morphological parameter is determined as the entity area.
[0112] In some embodiments, the morphological parameter value ranges of the above-mentioned vector entities are used to describe the geometric shape of the vector entities. The morphological parameter value ranges corresponding to vector entities of the same geometric shape are the same, and the morphological parameter value ranges corresponding to vector entities of different geometric shapes are different. The morphological parameter value ranges of the above-mentioned vector entities can be pre-set according to vector entities of different geometric shapes.
[0113] As an example, the morphological parameter value range of vector entity Y1 is U1, the morphological parameter value range of vector entity Y2 is U2, the value of the regional morphological parameter of reference entity area C1 is T1, and T1∈U1, the value of the regional morphological parameter of reference entity area C2 is T2, and T2∈U2, the value of the regional morphological parameter of reference entity area C1 is compared with the morphological parameter value range U1 and the morphological parameter value range U2, and the parameter comparison result corresponding to the reference entity area C1 is obtained. When the parameter comparison result corresponding to the reference entity area C1 indicates that the value of the regional morphological parameter is within the morphological parameter value range U1, the reference entity area C1 corresponding to the regional morphological parameter is determined to be a physical area; the value of the regional morphological parameter of the reference entity area C2 is compared with the morphological parameter value range U1 and the morphological parameter value range U2, and the parameter comparison result corresponding to the reference entity area C2 is obtained. When the parameter comparison result corresponding to the reference entity area C2 indicates that the value of the regional morphological parameter is within the morphological parameter value range U2, the reference entity area C2 corresponding to the regional morphological parameter is determined to be a physical area.
[0114] As an example, the morphological parameter value range of the vector entity Y1 is U1, the morphological parameter value range of the vector entity Y2 is U2, the value of the morphological parameter of the reference entity region C1 is T1, and T1 does not belong to U1 and U2, the value of the morphological parameter of the reference entity region C2 is T2, and T2 does not belong to U1 and U2, the value of the morphological parameter of the reference entity region C1 is compared with the morphological parameter value range U1 and the morphological parameter value range U2, and the parameter comparison result corresponding to the reference entity region C1 is obtained. The parameter comparison result corresponding to the reference entity region C1 indicates the value of the regional morphological parameter. When the value is not within the morphological parameter value range U1 and the morphological parameter value range U2, the reference entity area C1 corresponding to the regional morphological parameter is not determined as the entity area; the value of the regional morphological parameter of the reference entity area C2 is compared with the morphological parameter value range U1 and the morphological parameter value range U2 to obtain the parameter comparison result corresponding to the reference entity area C2. When the parameter comparison result corresponding to the reference entity area C2 indicates that the value of the regional morphological parameter is not within the morphological parameter value range U2 and the morphological parameter value range U1, the reference entity area C2 corresponding to the regional morphological parameter is not determined as the entity area.
[0115] In this way, by clustering the target pixel points in the vector map to be identified, at least one initial entity area in the vector map to be identified is obtained. When the number of initial entity areas is multiple, the initial entity areas in the vector map to be identified are clustered to obtain at least one reference entity area in the vector map to be identified. When the number of reference entity areas is multiple, the values of the regional morphological parameters of each reference entity area and the morphological parameter value range of the vector entity are obtained; the values of the morphological parameters of each region are compared with the morphological parameter value range respectively to obtain the parameter comparison results of the morphological parameters of each region; when the reference comparison result indicates that the value of the regional morphological parameter is within the morphological parameter value range, the reference entity area corresponding to the regional morphological parameter is determined as the entity area, so that through continuous clustering processing and comparison of the values of the regional morphological parameters with the morphological parameter value range, the entity area including at least one vector entity is accurately determined, thereby effectively improving the accuracy of the determined entity area.
[0116] In step 103, element recognition is performed on each vector element in the vector image to be recognized, and the entity category corresponding to each vector element is obtained.
[0117] In some embodiments, the element identification is a process of determining the entity category to which the vector element belongs.
[0118] In some embodiments, the above-mentioned element recognition is implemented through an element recognition network, which includes a recognition layer, a feature extraction layer and a classification layer.
[0119] As an example, see Figure 6 , Figure 6 is a schematic diagram of the principle of the element recognition network provided in the embodiment of the present application, Figure 6 The element recognition network shown includes a recognition layer 61 , a feature extraction layer 62 and a classification layer 63 .
[0120] In some embodiments, see Figure 7 , Figure 7 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 3 , Figure 3 The illustrated step 103 may be performed by Figure 7 Steps 1031 to 1033 are shown to be implemented.
[0121] In step 1031, the recognition layer is called to perform vector recognition on the vector image to be recognized, and obtain each vector element in the vector image to be recognized.
[0122] As an example, see Figure 6 , calling recognition layer 61, treating the recognition vector graph ( Figure 6 The above recognition layer is used to recognize the vector elements in the vector diagram to be recognized.
[0123] In step 1032, the feature extraction layer is called to perform feature extraction on each vector element to obtain element features corresponding to each vector element.
[0124] In some embodiments, the above element features correspond one-to-one to the vector elements, and the element features are vector expressions of the corresponding vector elements.
[0125] As an example, see Figure 6 , calling the feature extraction layer 62 to perform feature extraction on each vector element 611 respectively, and obtaining the element features corresponding to each vector element 611 respectively.
[0126] In step 1033, for each element feature, the classification layer is called, and based on the element feature, the corresponding vector element is classified to obtain the entity category corresponding to the vector element.
[0127] In some embodiments, the classification layer is used to classify vector elements to obtain entity categories corresponding to the vector elements.
[0128] As an example, see Figure 6 , for each element feature, the classification layer is called, and based on the element feature, the corresponding vector element 611 is classified to obtain the entity category corresponding to the vector element.
[0129] In some embodiments, the above step 1033 can be implemented as follows: for each feature element, call the classification layer, and based on the element characteristics, perform classification prediction on the corresponding vector element to obtain the probability that the vector element corresponds to each candidate entity category, and determine the candidate entity category with the highest probability as the entity category corresponding to the vector element.
[0130] In this way, by calling the recognition layer, vector recognition is performed on the vector image to be recognized, and each vector element in the vector image to be recognized is obtained. The feature extraction layer is called to extract features of each vector element respectively to obtain the element features corresponding to each vector element. For each element feature, the classification layer is called, and the corresponding vector elements are classified based on the element features to obtain the entity category corresponding to the vector element. In this way, the entity category corresponding to each vector element is accurately determined through the element recognition network, which effectively improves the accuracy of the entity category.
[0131] In step 104, entity recognition is performed on the vector image to be recognized in combination with the entity region and the entity category corresponding to each vector element to obtain an entity recognition result.
[0132] In some embodiments, the entity recognition result is used to indicate the entity category corresponding to each vector entity in the vector map to be recognized. The entity category corresponding to the vector entity is the same as the entity category corresponding to at least one vector element included in the vector entity.
[0133] In some embodiments, see Figure 8 , Figure 8 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 4 , Figure 3 Step 104 shown may be performed by Figure 8 Steps 1041A to 1043A are implemented as shown.
[0134] In step 1041A, the region confidence of each entity region and the element confidence of each vector element are obtained.
[0135] In some embodiments, the element confidence is used to indicate the probability that the vector element is a corresponding entity category, and the region confidence is used to indicate the probability that the number of vector entities in the entity region is one.
[0136] In some embodiments, the above-mentioned regional confidence is positively correlated with the prediction accuracy of the above-mentioned pixel recognition network, and the above-mentioned element confidence is positively correlated with the prediction accuracy of the above-mentioned element recognition network, that is, the greater the prediction accuracy of the above-mentioned pixel recognition network, the greater the corresponding regional confidence, and the smaller the prediction accuracy of the above-mentioned pixel recognition network, the smaller the corresponding regional confidence. The greater the prediction accuracy of the above-mentioned element recognition network, the greater the corresponding element confidence, and the smaller the prediction accuracy of the above-mentioned element recognition network, the smaller the corresponding element confidence.
[0137] In step 1042A, the confidences of the regions are summed to obtain a first confidence, and the confidences of the elements are summed to obtain a second confidence.
[0138] In some embodiments, the first confidence is positively correlated with the prediction accuracy of the pixel recognition network, and the second confidence is positively correlated with the prediction accuracy of the element recognition network, that is, the greater the prediction accuracy of the pixel recognition network, the greater the corresponding first confidence, and the smaller the prediction accuracy of the pixel recognition network, the smaller the corresponding first confidence. The greater the prediction accuracy of the element recognition network, the greater the corresponding second confidence, and the smaller the prediction accuracy of the element recognition network, the smaller the corresponding second confidence.
[0139] As an example, the expression of the first confidence level may be:
[0140] T=Q 1 +…+Q i +…+Q N (1)
[0141] Among them, T is used to indicate the first confidence level, Q 1 …Q i …Q N It is used to indicate the confidence level of each region, and N is used to indicate the number of regional confidence levels.
[0142] As an example, the expression of the second confidence level may be:
[0143] W=Y 1 +…+Y i +…+Y M (2)
[0144] Among them, W is used to indicate the second confidence level, Y 1 ,…Y i ,…Y M It is used to indicate the confidence of each element, and M is used to indicate the number of element confidences.
[0145] In step 1043A, entity recognition is performed on the vector image to be recognized in combination with the first confidence level and the second confidence level to obtain an entity recognition result.
[0146] In some embodiments, the above step 1043A can be implemented as follows: compare the first confidence level with the first confidence level threshold to obtain a first comparison result, and compare the second confidence level with the second confidence level threshold to obtain a second comparison result; and determine the entity recognition result by combining the first comparison result and the second comparison result.
[0147] In some embodiments, the first comparison result is used to indicate whether the first confidence level is greater than or equal to a first confidence threshold, and the second comparison result is used to indicate whether the second confidence level is greater than or equal to a second confidence threshold. The first confidence threshold is a critical confidence value for believing the first confidence level, and the second confidence threshold is a critical confidence value for believing the second confidence level.
[0148] In some embodiments, the entity recognition result is determined in combination with the first comparison result and the second comparison result, which can be achieved in the following manner: when the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, entity recognition is performed on the vector map to be recognized based on the entity category and entity area corresponding to the vector element to obtain an entity recognition result; when the first comparison result indicates that the first confidence is less than the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, entity recognition is performed on the vector map to be recognized based on the entity category corresponding to each vector element to obtain an entity recognition result; when the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is less than the second confidence threshold, entity recognition is performed on the vector map to be recognized based on the entity area to obtain an entity recognition result.
[0149] In some embodiments, when the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network and the prediction accuracy of the above-mentioned element recognition network are both high, and the obtained entity areas and entity categories of each vector element are relatively accurate. Then, based on the entity categories and entity areas corresponding to the vector elements, entity recognition can be performed on the vector image to be identified to obtain an entity recognition result, so that the obtained entity recognition result is more accurate.
[0150] In some embodiments, when the first comparison result indicates that the first confidence is less than the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network is low, while the prediction accuracy of the above-mentioned element recognition network is high. Then, entity recognition can be performed on the vector map to be identified based on the entity category corresponding to each vector element alone instead of the entity area to obtain an entity recognition result, thereby making the obtained entity recognition result more accurate.
[0151] In some embodiments, when the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is less than the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network is higher, while the prediction accuracy of the above-mentioned element recognition network is lower. In this case, instead of using the entity categories corresponding to each vector element for identification, entity recognition can be performed on the vector image to be identified based solely on the entity area to obtain an entity recognition result, thereby making the obtained entity recognition result more accurate.
[0152] In some embodiments, when the first comparison result indicates that the first confidence is less than the first confidence threshold, and the second comparison result indicates that the second confidence is less than the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network is low, and the prediction accuracy of the above-mentioned element recognition network is low. Then, the above-mentioned pixel recognition network and element recognition network can be further trained to improve the prediction accuracy of the pixel recognition network and the element recognition network. The confidences of the trained pixel recognition network and the element recognition network are compared again until the prediction accuracy of the pixel recognition network and the element recognition network meets the requirements, and then entity recognition is performed on the vector image to be recognized to obtain the entity recognition result, so that the obtained entity recognition result is more accurate.
[0153] In some embodiments, the above-mentioned entity recognition results include entity recognition results corresponding to each entity area respectively. The above-mentioned entity recognition of the vector image to be recognized based on the entity category and entity area corresponding to the vector element to obtain the entity recognition result can be achieved in the following way: the following processing is performed for each entity area respectively: the vector element located in the entity area is determined as the target vector element corresponding to the entity area; when the number of target vector elements is one, the entity category corresponding to the target vector element is determined as the entity recognition result corresponding to the entity area; when the number of target vector elements is multiple, the target vector elements are clustered based on the entity categories corresponding to the target vector elements to obtain at least one target vector entity in the entity area; the entity category of the target vector element in each target vector entity is determined as the entity category of the corresponding target vector entity; the entity category of each target vector entity is determined as the entity recognition result corresponding to the entity area.
[0154] In some embodiments, the target vector elements in the target vector entity respectively correspond to the same entity category.
[0155] In some embodiments, when the number of vector elements located in the entity region is one, it means that the number of target vector elements is one, and the entity category corresponding to the target vector element can be directly determined as the entity recognition result corresponding to the entity region.
[0156] As an example, when the number of vector elements located in the entity region R is one, the vector element in the entity region R is the target vector element R1, then the entity category corresponding to the target vector element R1 can be directly determined as the entity recognition result corresponding to the entity region R. The entity recognition result corresponding to the entity region R is used to indicate the entity category corresponding to the vector entity in the entity region R (the vector entity includes the target vector element R1).
[0157] In some embodiments, when the number of vector elements located in the entity area is multiple, it means that the number of target vector elements is multiple, and it is unknown how many target vector entities these multiple target vector elements are combined into. In this case, the target vector elements can be clustered based on the entity categories corresponding to the target vector elements to obtain at least one target vector entity in the entity area. The target vector elements in the target vector entity respectively correspond to the same entity category, and the entity category of the target vector elements in each target vector entity is determined as the entity category of the corresponding target vector entity, and the entity category of each target vector entity is determined as the entity recognition result corresponding to the entity area. The entity recognition result corresponding to the entity area is used to indicate the entity category of each target vector entity in the entity area.
[0158] As an example, when the number of vector elements located in the entity region R is multiple, the vector elements in the entity region R are target vector element R1, target vector element R2, and target vector element R3, the entity categories corresponding to target vector element R1 and target vector element R2 are the same, and target vector element R3 is different from the entity categories corresponding to target vector element R1 and target vector element R2. Then, based on the entity categories corresponding to the target vector elements, the target vector elements are clustered to obtain target vector entities P1 and target vector entities P2 in the entity region, the entity categories corresponding to the target vector elements R1 and target vector elements R2 in the target vector entity P1 are the same, and the target vector entity P2 includes the target vector element R3.
[0159] In this way, when the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network and the prediction accuracy of the above-mentioned element recognition network are both high. Then, at this time, for each entity area, the entity category of the vector elements in each entity area can be combined to determine the vector recognition results corresponding to each entity area, so that on the basis of the entity area, the entity category of the vector elements in the entity area can be further combined to generate the vector recognition results, so that the obtained vector recognition results are more accurate.
[0160] In some embodiments, the above-mentioned entity recognition of the vector graph to be identified based on the entity categories respectively corresponding to each vector element to obtain the entity recognition result can be achieved in the following way: when the number of vector elements in the vector graph to be identified is one, the entity category corresponding to the vector element is determined as the entity recognition result; when the number of vector elements in the vector graph to be identified is multiple, based on the entity categories respectively corresponding to the vector elements, the vector elements in the vector graph to be identified are clustered to obtain at least one vector entity in the vector graph to be identified; the entity categories of the vector elements in each vector entity are respectively determined as the entity categories of the corresponding vector entities; the entity category of each vector entity is determined as the entity recognition result.
[0161] In some embodiments, the entity categories corresponding to the vector elements in the vector entity are the same.
[0162] In some embodiments, when the number of vector elements in the vector image to be identified is one, it means that this one vector element constitutes a vector entity in the vector image to be identified, and the entity category corresponding to the vector element is determined as the entity recognition result corresponding to the vector image to be identified, and the entity recognition result corresponding to the vector image to be identified is used to indicate the entity category corresponding to this one vector entity in the vector image to be identified.
[0163] In some embodiments, when there are multiple vector elements in the vector graph to be identified, it means that it is unknown how many vector entities these multiple vector elements are combined into. In this case, the vector elements in the vector graph to be identified can be clustered based on the entity categories corresponding to the vector elements to obtain at least one vector entity in the vector graph to be identified. The entity categories corresponding to the vector elements in the vector entity are the same, and the entity categories of the vector elements in each vector entity are respectively determined as the entity categories of the corresponding vector entities; the entity categories of each vector entity are determined as entity recognition results, and the entity recognition results are used to indicate the entity categories of each vector entity in the vector graph to be identified.
[0164] As an example, when the number of vector elements in the vector map U to be identified is multiple, the vector elements in the vector map U to be identified are vector element T1, vector element T2 and vector element T3, vector element T1 and vector element T2 respectively correspond to the same entity category, and vector element T3 is different from the entity category corresponding to vector element T1 and vector element T2. Then, based on the entity categories corresponding to the vector elements, the vector elements are clustered to obtain vector entities J1 and vector entities J2 in the vector map U to be identified, the entity categories corresponding to vector elements T1 and vector elements T2 in vector entity J1 are the same, and the target vector entity J2 includes vector element T3.
[0165] In this way, when the first comparison result indicates that the first confidence is less than the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network is low, while the prediction accuracy of the above-mentioned element recognition network is high. Then, entity recognition can be performed on the vector image to be identified based on the entity categories corresponding to each vector element, and the vector elements in the vector image to be identified are clustered to obtain at least one vector entity in the vector image to be identified. The entity categories corresponding to the vector elements in the vector entity are the same, and the entity categories of the vector elements in each vector entity are respectively determined as the entity categories of the corresponding vector entities; the entity categories of each vector entity are determined as the entity recognition results, thereby effectively avoiding entity areas with low accuracy, thereby effectively reducing the influence of entity areas on the final entity recognition results, and making the obtained vector recognition results more accurate.
[0166] In some embodiments, the entity recognition result includes the entity recognition results corresponding to each entity area respectively. The above-mentioned entity recognition is performed on the vector image to be recognized based on the entity area to obtain the entity recognition result, which can be achieved in the following way: the following processing is performed for each entity area respectively: the vector elements located in the entity area are determined as the target vector elements corresponding to the entity area; when the number of target vector elements is one, the entity category corresponding to the target vector element is determined as the entity recognition result corresponding to the entity area; when the number of target vector elements is multiple, the expected entity contour is obtained, and based on the expected entity contour, the target vector elements are clustered to obtain at least one target vector entity in the entity area; the entity category corresponding to the expected entity contour is determined as the entity recognition result corresponding to the entity area.
[0167] In some embodiments, the contour formed by each target vector element in the target vector entity satisfies the desired entity contour.
[0168] In some embodiments, when the number of the target vector element is one, it means that the number of the target vector element is one, and then the entity category corresponding to the target vector element can be directly determined as the entity recognition result corresponding to the entity region.
[0169] As an example, when the number of vector elements located in the entity region R is one, the vector element in the entity region R is the target vector element R1, then the entity category corresponding to the target vector element R1 can be directly determined as the entity recognition result corresponding to the entity region R. The entity recognition result corresponding to the entity region R is used to indicate the entity category corresponding to the vector entity in the entity region R (the vector entity includes the target vector element R1).
[0170] In some embodiments, when there are multiple target vector elements, and it is unknown how many target vector entities these multiple target vector elements are combined into, since the accuracy of the entity categories corresponding to the target vector elements is low at this time, it is impossible to perform clustering processing according to the entity categories corresponding to the vector elements. In this case, the target vector elements can be clustered based on the expected entity contour to obtain at least one target vector entity in the entity area. The contour formed by each target vector element in the target vector entity satisfies the expected entity contour, and the entity category corresponding to the expected entity contour is determined as the entity recognition result corresponding to the entity area.
[0171] In this way, when there are multiple target vector elements, and it is unknown how many target vector entities these multiple target vector elements are combined into, since the accuracy of the entity categories corresponding to the target vector elements is low at this time, it is impossible to perform clustering processing through the entity categories corresponding to the vector elements. Then, the target vector elements can be clustered based on the expected entity contour to obtain at least one target vector entity in the entity area, and the contour formed by each target vector element in the target vector entity meets the expected entity contour, and the entity category corresponding to the expected entity contour is determined as the entity recognition result corresponding to the entity area. In this way, the entity categories corresponding to the target vector elements with low accuracy are effectively avoided, thereby effectively reducing the influence of the entity categories corresponding to the target vector elements on the final entity recognition result, so that the obtained vector recognition result is more accurate.
[0172] In some embodiments, see Fig. 9 , Fig. 9 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 5 , when the pixel recognition network and the element recognition network are fully trained and have high prediction accuracy, the entity recognition can be performed directly on the vector image to be identified based on the entity category and entity area corresponding to the vector elements to obtain the entity recognition result, which is explained below. Figure 3 The step 104 shown can be performed for each entity area separately. Fig. 9 Steps 1041B to 1045B are implemented as shown.
[0173] In step 1041B, the vector element located in the entity region is determined as the target vector element corresponding to the entity region.
[0174] As an example, the entity area L includes a vector element L1 , a vector element L2 , and a vector element L3 . Then, the vector element L1 , the vector element L2 , and the vector element L3 may be determined as the target vector elements corresponding to the entity area L .
[0175] In step 1042B, when the number of the target vector element is one, the entity category corresponding to the target vector element is determined as the entity recognition result corresponding to the entity region.
[0176] In some embodiments, when the number of vector elements located in the entity region is one, it means that the number of target vector elements is one, and the entity category corresponding to the target vector element can be directly determined as the entity recognition result corresponding to the entity region.
[0177] As an example, when the number of vector elements located in the entity region R is one, the vector element in the entity region R is the target vector element R1, then the entity category corresponding to the target vector element R1 can be directly determined as the entity recognition result corresponding to the entity region R. The entity recognition result corresponding to the entity region R is used to indicate the entity category corresponding to the vector entity in the entity region R (the vector entity includes the target vector element R1).
[0178] In step 1043B, when there are multiple target vector elements, clustering is performed on the target vector elements based on the entity categories corresponding to the target vector elements to obtain at least one target vector entity in the entity area.
[0179] In some embodiments, the target vector elements in the target vector entity respectively correspond to the same entity category.
[0180] In some embodiments, when the number of vector elements located in the entity area is multiple, it means that the number of target vector elements is multiple, and it is unknown how many target vector entities these multiple target vector elements are combined into. In this case, the target vector elements can be clustered based on the entity categories corresponding to the target vector elements to obtain at least one target vector entity in the entity area. The target vector elements in the target vector entity respectively correspond to the same entity category, and the entity category of the target vector elements in each target vector entity is determined as the entity category of the corresponding target vector entity, and the entity category of each target vector entity is determined as the entity recognition result corresponding to the entity area. The entity recognition result corresponding to the entity area is used to indicate the entity category of each target vector entity in the entity area.
[0181] As an example, when the number of vector elements located in the entity region R is multiple, the vector elements in the entity region R are target vector element R1, target vector element R2, and target vector element R3, the entity categories corresponding to target vector element R1 and target vector element R2 are the same, and target vector element R3 is different from the entity categories corresponding to target vector element R1 and target vector element R2. Then, based on the entity categories corresponding to the target vector elements, the target vector elements are clustered to obtain target vector entities P1 and target vector entities P2 in the entity region, the entity categories corresponding to the target vector elements R1 and target vector elements R2 in the target vector entity P1 are the same, and the target vector entity P2 includes the target vector element R3.
[0182] In step 1044B, the entity categories of the target vector elements in each target vector entity are respectively determined as the entity categories of the corresponding target vector entities.
[0183] Continuing with the above example, the entity category of the target vector element R3 in the target vector entity P2 is determined as the entity category of the target vector entity P2, and the entity categories of the target vector elements R1 and R2 in the target vector entity P1 are determined as the entity category corresponding to the target vector entity P1.
[0184] In step 1045B, the entity category of each target vector entity is determined as the entity recognition result corresponding to the entity region.
[0185] Continuing with the above example, the entity category corresponding to the target vector entity P1 and the entity category corresponding to the target vector entity P2 are determined as entity recognition results corresponding to the entity region.
[0186] Thus, by obtaining a vector image to be identified including at least one vector entity, performing image segmentation on the vector image to be identified, obtaining at least one entity region in the vector image to be identified, and performing element recognition on each vector element in the vector image to be identified, obtaining the entity category corresponding to each vector element, and combining the entity region and the entity category corresponding to the vector element, performing entity recognition on the vector image to be identified, and obtaining an entity recognition result. Thus, by performing image segmentation on the vector image to be identified, obtaining at least one entity region in the vector image to be identified, thereby identifying the region including the vector entity in the vector image to be identified from the overall macroscopic perspective, performing element recognition on each vector element in the vector image to be identified, obtaining the entity category corresponding to each vector element, thereby performing element recognition on the vector elements in the vector image to be identified from the dimensions of each vector element constituting the vector entity and from the detailed microscopic perspective, performing entity recognition on the vector image to be identified from the overall macroscopic entity region and the detailed microscopic entity category corresponding to each vector element, thereby making the obtained vector recognition result more accurate, and effectively improving the accuracy of entity recognition.
[0187] The following is an explanation of an exemplary application of the embodiments of the present application in an actual entity recognition application scenario.
[0188] When converting traditional architectural CAD data to 3D data in BIM, the traditional method is manual modeling, while AI can help to quickly extract CAD entities and convert them into BIM; traditional CAD AI extraction is generally based on image segmentation algorithms, that is, CAD is converted to an image during input, then AI is parsed, and then converted into an image recognition MASK, which is further combined with the original vector to extract vector entity information; however, this method is prone to line type recognition errors (lines are often intertwined).
[0189] The entity recognition method provided in the embodiment of the present application comprehensively utilizes vector AI extraction (selecting the base model based on the amount of training data) and image AI extraction (multiple graphic AI algorithms can be integrated, and different base models can be selected based on different amounts of data), and combined with business reality, business rules are integrated into the algorithm process, so that building BIM data that meets business needs can be extracted to support 3D generation. The vector + image transformer is used to innovatively identify the line business category of vector recognition and the pixel category of traditional image recognition, and the fusion analysis can effectively improve the accuracy of entity recognition compared to traditional image segmentation algorithms.
[0190] In some embodiments, see Fig.10 , Fig.10 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 6 , the entity recognition method provided in the embodiment of the present application can be Fig.10 Steps 201 to 209 are shown to be implemented.
[0191] In step 201, CAD software is used for marking.
[0192] In some embodiments, the CAD vector is annotated by CAD software to obtain first annotated data.
[0193] In step 202, labelme images are annotated.
[0194] In some embodiments, labelme image annotation annotates the CAD vector to obtain second annotation data.
[0195] In step 203, a space editor is called.
[0196] In some embodiments, the CAD vector is edited by a space editor to obtain edited data.
[0197] In step 204, SVG vector annotation data is generated.
[0198] In some embodiments, the first annotation data and the edited data are merged to obtain svg vector annotation data.
[0199] In step 205, image annotation json+image png is generated.
[0200] In some embodiments, the image annotation json+image png is generated based on the second annotation data.
[0201] In step 206, image annotation json+image png is generated.
[0202] In some embodiments, image annotation json+image png is generated based on the edit data.
[0203] In step 207, the vector image fusion AI model is called.
[0204] In step 208, TBIMS data is generated.
[0205] In step 209, the space editor 3D is loaded and rendered.
[0206] In some embodiments, see Fig.11 , Fig.11 It is a schematic diagram of the effect of the vector image to be identified provided in the embodiment of the present application. It is mainly through annotating CAD data based on multiple ways, such as vector type annotation based on the CAD environment, image annotation based on labelme, and image annotation based on the space editor to extract vector and image annotation data. The extracted annotation data is input into the AI model of vector and image fusion for vector entity recognition. The recognition result generates BIM data, and the three-dimensional display and rendering are performed through the space editor. For example, Fig.11 What is shown in the labelme image is a single door entity 71.
[0207] In some embodiments, graph neural network features can be incorporated into vector AI recognition, and traditional image features such as SIFT can also be added. In rule-based entity recognition, a vector set similarity algorithm based on machine learning is developed.
[0208] In some embodiments, see Fig.12 , Fig.12 This is a schematic diagram of the process of the entity recognition method provided in the embodiment of the present application. Figure 7 , the entity recognition method provided in the embodiment of the present application can be Fig.12 Steps 301 to 317 are implemented as shown.
[0209] In step 301, image data is acquired.
[0210] In step 302, image enhancement is performed.
[0211] In some embodiments, image enhancement is mainly performed based on albumenation, and mainly includes color enhancement, vertical flip enhancement, horizontal flip enhancement, random rotation of 90 degrees, edge enhancement, etc. Image enhancement is used to increase the number of image samples used to train the pixel recognition model.
[0212] In step 303, multiple CNN models are inferred separately.
[0213] In some embodiments, different CNN models are used for fusion according to different training data volumes and input image sizes. In the current application scenario, the input CAD image size is not fixed, so the image AI model should support image data input of unfixed size. If the training data volume is less than 10,000, the traditional image segmentation model is used, mainly Unet / Linknet / maskRCNN, etc., and the backbone uses efficientnet. If the training data volume is more than 10,000, MaskDino / maskFormer / DETR, etc. can also be used.
[0214] In step 304, multiple models are fused.
[0215] In some embodiments, multiple image segmentation models are fused. For semantic segmentation, model fusion can be performed by pixel weighted compounding. For instance segmentation, multiple models can be fused by credibility value comparison. Here, two models are sufficient.
[0216] In step 305, the image AI generates a physical region.
[0217] In some embodiments, for semantic segmentation, it is necessary to aggregate connected pixels to generate a solid area. The area value of the polygon and the aspect ratio of the outer envelope rectangle are calculated to exclude polygons that do not conform to the business type attributes. For instance segmentation, the area and aspect ratio of the obtained polygons are also counted to exclude interference data.
[0218] In step 306, the CAD data is vectorized.
[0219] In step 307, primitives are extracted from the vector CAD data.
[0220] In some embodiments, vector data extraction graphic primitive, tokenization module, Transformer module, and line semantic classification are implemented with reference to CAD Transformer. The CNN here considers that the transformer has a large demand for data volume. In this scenario of the embodiment of the present application, if the training data is less than 10,000, EfficientNet is used, and if it is greater than 10,000, Swin-L (IN21k) is used.
[0221] In step 308, feature extraction is performed on the primitives.
[0222] In step 309, the classification model is called.
[0223] In step 310 , line semantic categories are generated.
[0224] In step 311 , an original CAD vector is generated.
[0225] In step 312, the entity region and line semantic category data are combined for overlay analysis.
[0226] In some embodiments, if both are highly credible, the clustering method in CADTransformer is used for entity extraction, and the entity area generated by the image AI analysis is superimposed to obtain the overlapping vector entity. If the entity polygon generated by the image AI analysis has high credibility, and there are lines with low credibility in the line semantic category, the entity generated based on line clustering is not credible, and the polygon outline recognized by the image is more credible. At this time, vector entity extraction can be performed in a rule-based manner within the polygon.
[0227] In some embodiments, see Fig.13 , Fig.13 This is a schematic diagram of the principle of the entity area provided in an embodiment of the present application, taking a door as an example as follows: if the entity area boundary 91 is credible, the entity area boundary 91 is superimposed on the CAD vector map, and the lines with arcs are extracted as part of the door, and the two parallel lines adjacent to the arcs are taken as the other part of the door, and the two together constitute the door vector entity.
[0228] In some embodiments, if the credibility of the entity polygon recognized by the image AI is low, and the credibility of the line semantic category and line type is high, then the image recognition accuracy is often low and the contour is easily misidentified. For example, for types such as walls, it can be determined that the business type to be identified is vector entity extraction based on semantic clustering, and further filtering is performed based on rules, such as filtering whether it is a wall entity based on the aspect ratio of the entity's circumscribed rectangle.
[0229] The embodiment of the present application uses vector AI extraction (selecting the base model based on the amount of training data) and image AI extraction (which can integrate multiple graphic AI algorithms and select different base models based on different amounts of data), and combines business practice with business rules in the algorithm process, so that building BIM data that meets business needs can be extracted to support 3D generation. Compared with other solutions in the industry, the embodiment of the present application is more in line with business practice, and error transmission is targetedly controlled in the process, which can generate more accurate results data.
[0230] In step 313, when all the confidence levels are high, the lines in the region are semantically clustered to extract vector entities.
[0231] In step 314 , when the entity region has a high credibility, the line semantic category has a line type credibility lower than a threshold.
[0232] In step 315, when the entity region credibility is low and the line semantic category existence type credibility is higher than a threshold, semantic clustering is performed to extract vector entities.
[0233] In step 316, the rules are filtered.
[0234] In step 317, a vector entity is generated.
[0235] It is understandable that in the embodiments of the present application, related data such as the vector diagram to be identified are involved. When the embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0236] The following is a description of an exemplary structure of the entity identification device 455 provided in the embodiment of the present application implemented as a software module. In some embodiments, Figure 2 As shown, the software modules stored in the entity recognition device 455 of the memory 450 may include: an acquisition module, used to acquire a vector map to be recognized including at least one vector entity, wherein the vector entity includes at least one vector element; an image segmentation module, used to perform image segmentation on the vector map to be recognized, and obtain at least one entity region in the vector map to be recognized, wherein the entity region includes at least one of the vector entities; an element recognition module, used to perform element recognition on each of the vector elements in the vector map to be recognized, and obtain entity categories corresponding to each of the vector elements; an entity recognition module, used to perform entity recognition on the vector map to be recognized in combination with the entity region and the entity categories corresponding to each of the vector elements, and obtain an entity recognition result, wherein the entity recognition result is used to indicate the entity categories corresponding to each of the vector entities in the vector map to be recognized.
[0237] In some embodiments, the above-mentioned image segmentation module is also used to perform pixel point identification on each pixel point in the vector image to be identified, and obtain pixel identification results corresponding to each pixel point; when the pixel identification result indicates that the pixel point is a pixel point within the area where the vector element in the vector image to be identified is located, the pixel point is determined as a target pixel point; clustering processing is performed on the target pixel points in the vector image to be identified, and at least one initial entity area in the vector image to be identified is obtained, and the target pixel points in the initial entity area belong to the same vector element; based on the initial entity area, at least one entity area in the vector image to be identified is determined.
[0238] In some embodiments, the above-mentioned image segmentation module is also used to, when the number of the initial entity regions is one, determine the initial entity region as the entity region; when the number of the initial entity regions is multiple, perform clustering processing on the initial entity regions in the vector map to be identified to obtain at least one reference entity region in the vector map to be identified, and the initial entity regions in the reference entity regions belong to the same vector entity; based on the reference entity region, determine at least one of the entity regions in the vector map to be identified.
[0239] In some embodiments, the above-mentioned image segmentation module is also used to, when the number of the reference entity areas is one, determine the reference entity area as the entity area; when the number of the reference entity areas is multiple, obtain the value of the regional morphological parameters of each of the reference entity areas, and the value range of the morphological parameters of the vector entity; compare the value of each of the regional morphological parameters with the value range of the morphological parameter, respectively, to obtain the parameter comparison result of each of the regional morphological parameters; when the reference comparison result indicates that the value of the regional morphological parameter is within the value range of the morphological parameter, determine the reference entity area corresponding to the regional morphological parameter as the entity area.
[0240] In some embodiments, the above-mentioned element recognition is implemented through an element recognition network, and the element recognition network includes a recognition layer, a feature extraction layer and a classification layer. The above-mentioned element recognition module is also used to call the recognition layer to perform vector recognition on the vector image to be recognized, and obtain each of the vector elements in the vector image to be recognized; call the feature extraction layer to perform feature extraction on each of the vector elements respectively, and obtain the element features corresponding to each of the vector elements; for each of the element features, call the classification layer to classify the corresponding vector elements based on the element features, and obtain the entity category corresponding to the vector element.
[0241] In some embodiments, the above-mentioned entity recognition module is also used to obtain the regional confidence of each of the entity regions and the element confidence of each of the vector elements, the element confidence is used to indicate the probability that the vector element is the corresponding entity category, and the regional confidence is used to indicate the probability that the number of the vector entities in the entity region is one; summing up the regional confidences to obtain a first confidence, and summing up the element confidences to obtain a second confidence; combining the first confidence and the second confidence, performing entity recognition on the vector image to be identified to obtain the entity recognition result.
[0242] In some embodiments, the above-mentioned entity recognition module is also used to compare the first confidence and the first confidence threshold to obtain a first comparison result, and compare the second confidence and the second confidence threshold to obtain a second comparison result; when the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, based on the entity category corresponding to the vector element and the entity area, the entity recognition is performed on the vector image to be recognized to obtain the entity recognition result; when the first comparison result indicates that the first confidence is less than the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, based on the entity category corresponding to each of the vector elements, the entity recognition is performed on the vector image to be recognized to obtain the entity recognition result; when the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is less than the second confidence threshold, based on the entity area, the entity recognition is performed on the vector image to be recognized to obtain the entity recognition result.
[0243] In some embodiments, the entity recognition result includes entity recognition results corresponding to each of the entity regions, and the above-mentioned entity recognition module is also used to perform the following processing for each of the entity regions: determining the vector element located in the entity region as the target vector element corresponding to the entity region; when the number of the target vector element is one, determining the entity category corresponding to the target vector element as the entity recognition result corresponding to the entity region; when the number of the target vector elements is multiple, clustering the target vector elements based on the entity categories corresponding to the target vector elements, and obtaining at least one target vector entity in the entity region, and the entity categories corresponding to each of the target vector elements in the target vector entity are the same; determining the entity category of the target vector element in each of the target vector entities as the entity category of the corresponding target vector entity; determining the entity category of each of the target vector entities as the entity recognition result corresponding to the entity region.
[0244] In some embodiments, the above-mentioned entity recognition module is also used to, when the number of the vector elements in the vector image to be recognized is one, determine the entity category corresponding to the vector element as the entity recognition result; when the number of the vector elements in the vector image to be recognized is multiple, based on the entity categories respectively corresponding to the vector elements, cluster the vector elements in the vector image to be recognized to obtain at least one vector entity in the vector image to be recognized, and the entity categories respectively corresponding to the vector elements in the vector entity are the same; determine the entity categories of the vector elements in each of the vector entities as the entity categories of the corresponding vector entities; and determine the entity categories of each of the vector entities as the entity recognition result.
[0245] In some embodiments, the entity recognition result includes entity recognition results corresponding to each of the entity regions, and the above-mentioned entity recognition module is also used to perform the following processing for each of the entity regions: determining the vector elements located in the entity region as the target vector elements corresponding to the entity region; when the number of the target vector elements is one, determining the entity category corresponding to the target vector element as the entity recognition result corresponding to the entity region; when the number of the target vector elements is multiple, obtaining the expected entity contour, and based on the expected entity contour, clustering the target vector elements to obtain at least one target vector entity in the entity region, and the contour formed by each of the target vector elements in the target vector entity satisfies the expected entity contour; determining the entity category corresponding to the expected entity contour as the entity recognition result corresponding to the entity region.
[0246] In some embodiments, the above-mentioned entity recognition module is also used to perform the following processing for each of the entity areas: determining the vector elements located in the entity area as the target vector elements corresponding to the entity area; when the number of the target vector elements is one, determining the entity category corresponding to the target vector element as the entity recognition result corresponding to the entity area; when the number of the target vector elements is multiple, clustering the target vector elements based on the entity categories respectively corresponding to the target vector elements to obtain at least one target vector entity in the entity area, and the entity categories respectively corresponding to the target vector elements in the target vector entity are the same; determining the entity categories of the target vector elements in each of the target vector entities as the entity categories of the corresponding target vector entities; determining the entity categories of each of the target vector entities as the entity recognition result corresponding to the entity area.
[0247] The embodiment of the present application provides a computer program product, which includes a computer program or a computer executable instruction, and the computer program or the computer executable instruction is stored in a computer-readable storage medium. The processor of the electronic device reads the computer executable instruction from the computer-readable storage medium, and the processor executes the computer executable instruction, so that the electronic device performs the entity recognition method described in the embodiment of the present application.
[0248] The embodiment of the present application provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor will be caused to execute the entity recognition method provided by the embodiment of the present application, for example, Figure 3 The entity recognition method shown.
[0249] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface storage, optical disk, or CD-ROM; or it may be various electronic devices including one or any combination of the above memories.
[0250] In some embodiments, computer executable instructions may be in the form of a program, software, software module, script or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a stand-alone program or as a module, component, subroutine or other unit suitable for use in a computing environment.
[0251] As an example, computer-executable instructions may, but need not, correspond to a file in a file system, may be stored as part of a file storing other programs or data, such as in one or more scripts in a HyperText Markup Language (HTML) document, in a single file dedicated to the program in question, or in multiple coordinated files (e.g., files storing one or more modules, subroutines, or code portions).
[0252] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0253] As an example, computer executable instructions may be deployed to be executed on one electronic device, or on multiple electronic devices located at one site, or on multiple electronic devices distributed at multiple sites and interconnected by a communication network.
[0254] In summary, the embodiments of the present application have the following beneficial effects:
[0255] (1) By obtaining a vector image to be identified including at least one vector entity, performing image segmentation on the vector image to be identified, obtaining at least one entity region in the vector image to be identified, and performing element recognition on each vector element in the vector image to be identified, obtaining the entity category corresponding to each vector element, and combining the entity region and the entity category corresponding to each vector element, performing entity recognition on the vector image to be identified, and obtaining an entity recognition result. In this way, by performing image segmentation on the vector image to be identified, obtaining at least one entity region in the vector image to be identified, the region including the vector entity in the vector image to be identified is identified from the overall macroscopic perspective, and by performing element recognition on each vector element in the vector image to be identified, the entity category corresponding to each vector element is obtained, thereby performing element recognition on the vector elements in the vector image to be identified from the dimensions of each vector element constituting the vector entity and from the detailed microscopic perspective, and by performing entity recognition on the vector image to be identified from the overall macroscopic entity region and the detailed microscopic entity category corresponding to each vector element, the vector image to be identified is made more accurate, thereby effectively improving the accuracy of entity recognition.
[0256] (2) By clustering the target pixel points in the vector image to be identified, at least one initial entity region in the vector image to be identified is obtained. When the number of initial entity regions is multiple, the initial entity regions in the vector image to be identified are clustered to obtain at least one reference entity region in the vector image to be identified. When the number of reference entity regions is multiple, the values of the regional morphological parameters of each reference entity region and the value range of the morphological parameters of the vector entity are obtained; the values of the morphological parameters of each region are compared with the value range of the morphological parameters respectively to obtain the parameter comparison results of the morphological parameters of each region; when the reference comparison result indicates that the value of the regional morphological parameter is within the value range of the morphological parameter, the reference entity region corresponding to the regional morphological parameter is determined as the entity region, so that through continuous clustering processing and comparison of the values of the regional morphological parameters with the value range of the morphological parameters, the entity region including at least one vector entity is accurately determined, thereby effectively improving the accuracy of the determined entity region.
[0257] (3) By calling the recognition layer, vector recognition is performed on the vector image to be recognized to obtain each vector element in the vector image to be recognized, and the feature extraction layer is called to extract features of each vector element to obtain the element features corresponding to each vector element. For each element feature, the classification layer is called to classify the corresponding vector elements based on the element features to obtain the entity category corresponding to the vector element, thereby accurately determining the entity category corresponding to each vector element through the element recognition network, effectively improving the accuracy of the entity category.
[0258] (4) When the first comparison result indicates that the first confidence is less than the first confidence threshold, and the second comparison result indicates that the second confidence is less than the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network is low, and the prediction accuracy of the above-mentioned element recognition network is low. Then, the above-mentioned pixel recognition network and element recognition network can be further trained to improve the prediction accuracy of the pixel recognition network and the element recognition network. The confidences of the trained pixel recognition network and the element recognition network are compared again until the prediction accuracy of the pixel recognition network and the element recognition network meets the requirements. Then, entity recognition is performed on the vector image to be recognized to obtain the entity recognition result, so that the obtained entity recognition result is more accurate.
[0259] (5) When the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is less than the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network is high, while the prediction accuracy of the above-mentioned element recognition network is low. In this case, instead of using the entity categories corresponding to each vector element for recognition, entity recognition can be performed on the vector image to be recognized based solely on the entity area to obtain an entity recognition result, thereby making the obtained entity recognition result more accurate.
[0260] (6) When the first comparison result indicates that the first confidence is less than the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network is low, while the prediction accuracy of the above-mentioned element recognition network is high. In this case, entity recognition can be performed on the vector map to be recognized based on the entity categories corresponding to each vector element instead of the entity area, thereby obtaining an entity recognition result, thereby making the obtained entity recognition result more accurate.
[0261] (7) When the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network and the prediction accuracy of the above-mentioned element recognition network are both high, and the obtained entity area and the entity category of each vector element are relatively accurate. Then, based on the entity category and entity area corresponding to the vector element, entity recognition can be performed on the vector image to be identified to obtain an entity recognition result, so that the obtained entity recognition result is relatively accurate.
[0262] (8) When the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network and the prediction accuracy of the above-mentioned element recognition network are both high. Then, for each entity area, the entity category of the vector elements in each entity area can be combined to determine the vector recognition results corresponding to each entity area. On the basis of the entity area, the entity category of the vector elements in the entity area can be further combined to generate the vector recognition results, so that the obtained vector recognition results are more accurate.
[0263] (9) When the first comparison result indicates that the first confidence is less than the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, it means that the prediction accuracy of the above-mentioned pixel recognition network is low, while the prediction accuracy of the above-mentioned element recognition network is high. Then, entity recognition can be performed on the vector image to be recognized based on the entity categories corresponding to each vector element alone, and the vector elements in the vector image to be recognized are clustered to obtain at least one vector entity in the vector image to be recognized. The entity categories corresponding to the vector elements in the vector entity are the same. The entity categories of the vector elements in each vector entity are respectively determined as the entity categories of the corresponding vector entities; the entity categories of each vector entity are determined as the entity recognition results, thereby effectively avoiding the entity areas with low accuracy, thereby effectively reducing the influence of the entity areas on the final entity recognition results, and making the obtained vector recognition results more accurate.
[0264] (10) When there are multiple target vector elements, and it is unknown how many target vector entities these multiple target vector elements are combined into, since the accuracy of the entity categories corresponding to the target vector elements is low at this time, it is impossible to perform clustering processing based on the entity categories corresponding to the vector elements. Then, the target vector elements can be clustered based on the expected entity contour to obtain at least one target vector entity in the entity area. The contour formed by each target vector element in the target vector entity satisfies the expected entity contour, and the entity category corresponding to the expected entity contour is determined as the entity recognition result corresponding to the entity area. In this way, the entity categories corresponding to the target vector elements with low accuracy are effectively avoided, thereby effectively reducing the influence of the entity categories corresponding to the target vector elements on the final entity recognition result, so that the obtained vector recognition result is more accurate.
[0265] The above is only an embodiment of the present application and is not intended to limit the protection scope of the present application. Any modifications, equivalent substitutions and improvements made within the spirit and scope of the present application are included in the protection scope of the present application.
Claims
1. A method for entity recognition, It is characterized in that The method comprises: Acquire a vector graph to be identified including at least one vector entity, wherein the vector entity includes at least one vector element; Performing image segmentation on the vector image to be identified to obtain at least one entity region in the vector image to be identified, wherein the entity region includes at least one of the vector entities; Performing element recognition on each of the vector elements in the vector image to be recognized to obtain entity categories corresponding to each of the vector elements; In combination with the entity area and the entity categories corresponding to each of the vector elements, entity recognition is performed on the vector image to be identified to obtain an entity recognition result, and the entity recognition result is used to indicate the entity categories corresponding to each of the vector entities in the vector image to be identified.
2. The method according to claim 1, It is characterized in that The performing image segmentation on the vector image to be identified to obtain at least one entity region in the vector image to be identified includes: Perform pixel recognition on each pixel in the vector image to be recognized, and obtain pixel recognition results corresponding to each pixel; When the pixel recognition result indicates that the pixel point is a pixel point within the region where the vector element in the vector graph to be recognized is located, determining the pixel point as a target pixel point; Performing clustering processing on the target pixel points in the vector image to be identified to obtain at least one initial entity region in the vector image to be identified, wherein the target pixel points in the initial entity region belong to the same vector element; Based on the initial entity region, at least one entity region in the vector map to be identified is determined.
3. The method according to claim 2, It is characterized in that The determining, based on the initial entity region, at least one entity region in the vector map to be identified comprises: When the number of the initial entity regions is one, determining the initial entity region as the entity region; When there are multiple initial entity regions, clustering the initial entity regions in the vector graph to be identified to obtain at least one reference entity region in the vector graph to be identified, wherein the initial entity regions in the reference entity regions belong to the same vector entity; At least one of the entity regions in the vector map to be identified is determined based on the reference entity region.
4. The method according to claim 3, It is characterized in that The determining, based on the reference entity region, at least one entity region in the vector map to be identified comprises: When the number of the reference entity regions is one, determining the reference entity region as the entity region; When there are multiple reference entity regions, obtaining the value of the region morphological parameter of each reference entity region and the value range of the morphological parameter of the vector entity; Comparing the values of the morphological parameters of each region with the value range of the morphological parameters respectively, to obtain parameter comparison results of the morphological parameters of each region; When the reference comparison result indicates that the value of the regional morphological parameter is within the morphological parameter value range, the reference entity region corresponding to the regional morphological parameter is determined as the entity region.
5. The method according to claim 1, It is characterized in that The element recognition is implemented by an element recognition network, which includes a recognition layer, a feature extraction layer and a classification layer. The element recognition is performed on each of the vector elements in the vector map to be recognized to obtain the entity category corresponding to each of the vector elements, including: Calling the recognition layer to perform vector recognition on the vector graph to be recognized, and obtaining each of the vector elements in the vector graph to be recognized; Calling the feature extraction layer to extract features from each of the vector elements to obtain element features corresponding to each of the vector elements; For each of the element features, the classification layer is called, and based on the element features, the corresponding vector elements are classified to obtain the entity category corresponding to the vector element.
6. The method according to claim 1, It is characterized in that Combining the entity region and the entity categories corresponding to each of the vector elements, performing entity recognition on the vector image to be recognized, and obtaining an entity recognition result, includes: Obtaining a region confidence of each of the entity regions and an element confidence of each of the vector elements, wherein the element confidence is used to indicate a probability that the vector element is a corresponding entity category, and the region confidence is used to indicate a probability that the number of the vector entities in the entity region is one; The confidences of the regions are summed to obtain a first confidence, and the confidences of the elements are summed to obtain a second confidence; The first confidence level and the second confidence level are combined to perform entity recognition on the vector image to be recognized to obtain the entity recognition result.
7. The method according to claim 6, It is characterized in that The combining the first confidence level and the second confidence level to perform entity recognition on the vector image to be recognized to obtain the entity recognition result includes: Comparing the first confidence level with a first confidence level threshold to obtain a first comparison result, and comparing the second confidence level with a second confidence level threshold to obtain a second comparison result; When the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, performing entity recognition on the vector map to be recognized based on the entity category and the entity area corresponding to the vector element to obtain the entity recognition result; When the first comparison result indicates that the first confidence is less than the first confidence threshold, and the second comparison result indicates that the second confidence is greater than or equal to the second confidence threshold, performing entity recognition on the vector map to be recognized based on the entity categories corresponding to the vector elements, to obtain the entity recognition result; When the first comparison result indicates that the first confidence is greater than or equal to the first confidence threshold, and the second comparison result indicates that the second confidence is less than the second confidence threshold, entity recognition is performed on the vector image to be identified based on the entity area to obtain the entity recognition result.
8. The method according to claim 7, It is characterized in that The entity recognition result includes entity recognition results corresponding to each of the entity regions. The entity recognition is performed on the vector image to be recognized based on the entity category and the entity region corresponding to the vector element to obtain the entity recognition result, including: The following processing is performed for each of the entity regions: Determine the vector element located in the entity area as the target vector element corresponding to the entity area; When the number of the target vector element is one, determining the entity category corresponding to the target vector element as the entity recognition result corresponding to the entity region; When there are multiple target vector elements, clustering the target vector elements based on the entity categories respectively corresponding to the target vector elements to obtain at least one target vector entity in the entity area, wherein the entity categories respectively corresponding to the target vector elements in the target vector entity are the same; Determine the entity category of the target vector element in each target vector entity as the entity category of the corresponding target vector entity; The entity category of each of the target vector entities is determined as the entity recognition result corresponding to the entity region.
9. The method according to claim 7, It is characterized in that The performing entity recognition on the vector image to be recognized based on the entity categories corresponding to the vector elements to obtain the entity recognition result includes: When the number of the vector elements in the vector graph to be identified is one, determining the entity category corresponding to the vector element as the entity recognition result; When there are multiple vector elements in the vector diagram to be identified, clustering the vector elements in the vector diagram to be identified based on the entity categories respectively corresponding to the vector elements to obtain at least one vector entity in the vector diagram to be identified, wherein the entity categories respectively corresponding to the vector elements in the vector entity are the same; Determine the entity category of the vector elements in each of the vector entities as the entity category of the corresponding vector entity; The entity category of each of the vector entities is determined as the entity recognition result.
10. The method according to claim 7, It is characterized in that The entity recognition result includes entity recognition results corresponding to each of the entity regions. Based on the entity regions, entity recognition is performed on the vector image to be recognized to obtain the entity recognition result, including: The following processing is performed for each of the entity regions: Determine the vector element located in the entity area as the target vector element corresponding to the entity area; When the number of the target vector element is one, determining the entity category corresponding to the target vector element as the entity recognition result corresponding to the entity region; When the number of the target vector elements is multiple, obtaining an expected entity contour, and based on the expected entity contour, performing clustering processing on the target vector elements to obtain at least one target vector entity in the entity area, wherein the contour formed by each of the target vector elements in the target vector entity satisfies the expected entity contour; The entity category corresponding to the expected entity contour is determined as the entity recognition result corresponding to the entity area.
11. The method according to claim 1, It is characterized in that Combining the entity region and the entity categories corresponding to each of the vector elements, performing entity recognition on the vector image to be recognized, and obtaining an entity recognition result, includes: The following processing is performed for each of the entity regions: Determine the vector element located in the entity area as the target vector element corresponding to the entity area; When the number of the target vector element is one, determining the entity category corresponding to the target vector element as the entity recognition result corresponding to the entity region; When there are multiple target vector elements, clustering the target vector elements based on the entity categories respectively corresponding to the target vector elements to obtain at least one target vector entity in the entity area, wherein the entity categories respectively corresponding to the target vector elements in the target vector entity are the same; Determine the entity category of the target vector element in each target vector entity as the entity category of the corresponding target vector entity; The entity category of each of the target vector entities is determined as the entity recognition result corresponding to the entity region.
12. An entity recognition device, It is characterized in that The device comprises: An acquisition module, used for acquiring a vector graph to be identified including at least one vector entity, wherein the vector entity includes at least one vector element; An image segmentation module, configured to perform image segmentation on the vector image to be identified, and obtain at least one entity region in the vector image to be identified, wherein the entity region includes at least one of the vector entities; An element recognition module is used to perform element recognition on each of the vector elements in the vector image to be recognized, and obtain entity categories corresponding to each of the vector elements; The entity recognition module is used to perform entity recognition on the vector map to be recognized in combination with the entity area and the entity categories corresponding to each of the vector elements to obtain an entity recognition result, wherein the entity recognition result is used to indicate the entity categories corresponding to each of the vector entities in the vector map to be recognized.
13. An electronic device, It is characterized in that The electronic device comprises: A memory for storing computer executable instructions or computer programs; A processor, configured to implement the entity recognition method according to any one of claims 1 to 11 when executing the computer executable instructions or computer programs stored in the memory.
14. A computer-readable storage medium storing computer-executable instructions, It is characterized in that When the computer executable instructions are executed by a processor, the entity recognition method according to any one of claims 1 to 11 is implemented.
15. A computer program product comprising a computer program or computer executable instructions, It is characterized in that When the computer program or computer executable instructions are executed by a processor, the entity recognition method according to any one of claims 1 to 11 is implemented.