Method, apparatus, device, and medium for obtaining industry category of POI
By combining multimodal information of signboard images and door face images, a cross-modal graphic and text recall model is used to obtain the first-level industry category of POI, which solves the problem of insufficient recall and accuracy in the existing technology, and achieves high recall and high accuracy in the POI industry category.
Patent Information
- Application Number
- CN202210161845.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-22
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-02-22
AI Technical Summary
The industry categories in which prior art are difficult to obtain POIs accurately and efficiently, especially when there are dirty tags in the POI names or the names are not obvious, resulting in insufficient recall and accuracy.
Combining the multimodal information of the POI's signature image and the door face image, using the pre-trained cross-modal graphic and text recall model, the first-level industry categories of POI are obtained through the multimodal features of image features and text features, including the training of the double tower model and the scoring of the industry category.
It improves the recall and accuracy of the POI industry category, and can clean the door face image information in the presence of dirty tags, ensuring high recall and high accuracy of the industry category.
Smart Images

Figure CN115205612B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technologies, and in particular to intelligent transportation, NLP, and deep learning technologies. Specifically, it relates to a method, apparatus, device, medium, and program product for obtaining the industry category of a POI. Background Art
[0002] POI (Point Of Interest) generally refers to individual entity points on a map. POI contains many attributes. In addition to the name and address, there is also an important attribute that needs to be constructed: the industry category label Tag. Tag generally refers to the industry category of this POI. According to the latest classification standard, the categories of Tag are divided into primary and secondary levels, presenting a tree-like structure. For example, the secondary Tags of the primary Tag "Food" can be: Chinese restaurant, foreign restaurant, or coffee shop, etc.
[0003] Tag is very important reference information in downstream applications such as recommendation and retrieval. Therefore, accurately recalling the Tag of a POI is of great significance. Summary of the Invention
[0004] The present disclosure provides a method, apparatus, device, medium, and program product for obtaining the industry category of a POI.
[0005] According to one aspect of the present disclosure, a method for obtaining the industry category of a POI is provided, including:
[0006] Obtaining the signboard image and the facade image of the POI;
[0007] Using a pre-trained cross-modal image-text retrieval model, and outputting the target primary industry category of the POI according to the signboard image and the facade image.
[0008] According to another aspect of the present disclosure, an apparatus for obtaining the industry category of a POI is provided, including:
[0009] An image acquisition module, configured to obtain the signboard image and the facade image of the POI;
[0010] An industry category acquisition module, configured to use a pre-trained cross-modal image-text retrieval model, and output the target primary industry category of the POI according to the signboard image and the facade image.
[0011] According to another aspect of the present disclosure, an electronic device is provided, including:
[0012] At least one processor; and
[0013] A memory communicatively connected to the at least one processor; wherein,
[0014] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for obtaining the industry category of a POI according to any embodiment of the present disclosure.
[0015] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method for obtaining the industry category of a POI according to any embodiment of the present disclosure.
[0016] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method for obtaining the industry category of a POI according to any embodiment of the present disclosure.
[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0019] Figure 1 is a schematic diagram of a method for obtaining the industry category of a POI according to an embodiment of the present disclosure;
[0020] Figure 2 is a schematic diagram of a method for obtaining the industry category of a POI according to an embodiment of the present disclosure;
[0021] Figure 3a is a schematic diagram of a method for obtaining the industry category of a POI according to an embodiment of the present disclosure;
[0022] Figure 3b is a schematic diagram of a method for obtaining the industry category of a POI according to an embodiment of the present disclosure;
[0023] Figure 4 is a schematic structural diagram of a device for obtaining the industry category of a POI according to an embodiment of the present disclosure;
[0024] Figure 5 is a block diagram of an electronic device for implementing the method for obtaining the industry category of a POI according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0026] Figure 1 is a schematic flowchart of a method for obtaining the industry category of a POI according to an embodiment of the present disclosure. This embodiment is suitable for obtaining the industry category to which the POI belongs, relates to the field of image processing technology, and particularly relates to intelligent transportation, NLP, and deep learning technologies. This method can be executed by a device for obtaining the industry category of a POI, which is implemented in a software and / or hardware manner, preferably configured in an electronic device, such as a computer device or a server, etc. As Figure 1 shown, the method specifically includes the following:
[0027] S101. Obtain the sign image and the facade image of the POI.
[0028] For example, the sign image and the facade image are extracted from the original image of the POI using image recognition technology. Among them, the facade image usually refers to the image of the entrance of the POI and a certain area around it. The sign can be located above the facade, or on the left or right side of the facade.
[0029] S102. Use a pre-trained cross-modal image-text retrieval model to output the target primary industry category of the POI according to the sign image and the facade image.
[0030] The signboard image usually contains text information such as the name of the POI. By using technologies such as OCR (Optical Character Recognition), the text information can be extracted from the signboard image. The facade image usually contains image information of the POI facade and its surrounding environment. This environmental image information has a certain correlation with the industry category of the POI. For example, around the facade of a POI in the "beauty" industry, there may be portraits of women. Around the facade of a "western restaurant" in the "food" industry, there may be images of western food dishes. Therefore, in addition to combining the text information in the signboard image, the embodiments of the present disclosure also combine the image information in the facade image, and use this multimodal information as the discrimination basis for the industry category together, so as to avoid the problem that when determining the industry category only based on the POI name in the signboard, the industry category cannot be accurately obtained due to the presence of dirty tags in the POI name, and can further improve the accuracy of the recalled industry category. At the same time, due to the diversity of the POI names in the signboard image, some names do not have obvious industry definitions, so the industry category cannot be obtained through semantic analysis of the POI name. However, when combining multimodal information as the discrimination basis for the industry category together, this problem can be overcome, and the recall rate of the industry category can be improved.
[0031] Among them, the cross-modal image-text recall model can be any cross-modal model in the prior art, and the signboard image and facade image of the POI in the existing database and the industry category corresponding to the POI are used as training samples, and obtained through model training. The embodiments of the present disclosure do not make any limitations on the network structure of the cross-modal model.
[0032] The technical solution of the embodiments of the present disclosure combines the multimodal information of the signboard image and facade image of the POI, and uses the cross-modal image-text recall model to obtain the primary industry category of the POI, which can not only ensure a high recall rate of the POI industry category, but also improve the accuracy of the recalled industry category.
[0033] Figure 2 It is a schematic flowchart of the method for obtaining the industry category of the POI according to the embodiments of the present disclosure. This embodiment is further optimized on the basis of the above embodiment. As Figure 2 shown, the method specifically includes the following:
[0034] S201. Obtain the signboard image and facade image of the POI.
[0035] For example, the POI original image containing the signboard image and facade image can be obtained from the existing database first, and then any image recognition method in the prior art can be used to extract the signboard image of the POI from the POI original image. The embodiments of the present disclosure do not make any limitations on the above image recognition method.
[0036] In addition, the facade image of the POI can be obtained in the following way: determine the coordinates of the signboard image on the original image of the POI; according to the coordinates, obtain the environmental image within a set range around the signboard image in the original image, and use the environmental image as the facade image of the POI, where the set range is determined according to the length and width of the signboard image.
[0037] Specifically, the types of signboard images can include horizontal signboard images and vertical signboard images. Horizontal signboard images are usually located above the facade of the POI, and vertical signboard images are usually located on the left or right side of the facade of the POI. The coordinates of the signboard image can include the coordinates of at least one vertex or endpoint of the signboard image. Based on these coordinates, the size of the signboard image can be determined, including the length and width, etc. When obtaining the facade image, starting from the coordinates of the signboard image, according to the type of the signboard image, extend downward, to the right or to the left in the original image of the POI. The extended width and distance are related to the length and width of the signboard image. For example, both the extended width and distance are the same as the length of the signboard image. After extension, the facade image can be intercepted.
[0038] S202. Use the pre-trained cross-modal image-text retrieval model to score the industry categories in the industry category database according to the signboard image and the facade image.
[0039] The cross-modal image-text retrieval model is a two-tower model. The two-tower model includes an image tower for obtaining the image features in the signboard image and the facade image, and a text tower for obtaining the text features in the signboard image and the facade image. Thus, the industry categories are obtained through the multi-modal features of the image features and the text features.
[0040] Among them, the training process of the cross-modal image-text retrieval model includes: obtaining POI sample data, where the sample data includes the signboard image, the facade image and the corresponding industry category of the POI, and the sample data is divided into positive samples and negative samples; input the sample data into the pre-built cross-modal image-text retrieval model, and through the image tower and the text tower in the cross-modal image-text retrieval model, extract the image features and text features in the signboard image and the facade image respectively, and obtain the vector representations of the image features and the text features; during the training process, use the preset loss function to make the distance between the vector representations of the image features and the text features and the vector representation of the corresponding industry category in the positive samples tend to approach, and the distance between the vector representations of the image features and the text features and the vector representation of the corresponding industry category in the negative samples tend to be far away. Through the training of a large amount of sample data, the final cross-modal image-text retrieval model can be obtained.
[0041] The trained cross-modal image-text retrieval model can extract the image features and text features in the signboard image and the facade image of the POI input into the model, and then score each industry category based on the feature vectors of the image features and text features and the feature vectors of each industry category in the industry category database.
[0042] S203. According to the scoring results, output a candidate industry category set including a set number of candidate industry categories, where the candidate industry categories in the candidate industry category set include first-level categories and / or second-level categories.
[0043] Among them, a scoring threshold can be set in advance, and at least one industry category with a score exceeding the scoring threshold can be selected from each industry category according to the scoring results. Further, a set number of industry categories can be selected from at least one industry category in descending order of scores as candidate industry categories and form a candidate industry category set. For example, select the top 5 industry categories with the highest scores as candidate industry categories. And the candidate industry categories can be first-level categories, second-level categories, or both first-level categories and second-level categories.
[0044] S204. If, in the candidate industry category set, there is a set proportion of candidate industry categories corresponding to the same first-level category, then take this first-level category as the target first-level industry category of the POI.
[0045] Among them, the set proportion can be more than half. For example, among the top 5 candidate industry categories with the highest scores, if more than half (3) of the candidate industry categories correspond to the same first-level category, then this first-level category can be taken as the target first-level industry category of the POI. Exemplarily, if among the top 5 candidate industry categories, they are the first-level category "Food", the second-level category "Western Restaurant", the second-level category "Chinese Restaurant", the first-level category "Beauty" and the second-level category "Mall" respectively, then since the first-level categories corresponding to the first three industry categories are all "Food", the final target first-level industry category is "Food". Thus, even in the case of errors in the candidate industry categories, the final target first-level industry category can be more accurately identified.
[0046] The technical solution of the embodiment of the present disclosure combines the multi-modal information of the signboard image and the facade image of the POI, uses a cross-modal image-text recall model to score the industry categories in the industry category database, and selects a set of candidate industry categories according to the scores. Then, by determining whether there are the same first-level industry categories corresponding to a set proportion of the candidate industry categories in the set, the final target first-level industry category is determined, which can not only ensure a high recall rate of the POI industry category, but also improve the accuracy of the recalled industry category. For example, if there are dirty tags in the signboard of the POI, the facade image can be used to further clean them to improve the accuracy of the tags; in addition, for some POIs that cannot be labeled by the POI name, the facade image information can also be used to supplement and recall the first-level tags, thus significantly improving the tag recall rate of the POIs in the map.
[0047] Figure 3a FIG. is a schematic diagram of a method for obtaining the industry category of a POI according to an embodiment of the present disclosure. This embodiment is further optimized on the basis of the above embodiment. As Figure 3a shown, the method specifically includes the following:
[0048] S301. Obtain the signboard image and the facade image of the POI.
[0049] S302. Use a pre-trained cross-modal image-text recall model to output the target first-level industry category of the POI according to the signboard image and the facade image.
[0050] S303. Extract the text information in the signboard image.
[0051] For example, extract the text information in the signboard image through OCR technology.
[0052] S304. Obtain the POI name from the text information by performing layout analysis on the signboard image.
[0053] Among them, the purpose of layout analysis is to distinguish which of the text information belongs to the POI name and which belongs to other texts such as the business scope. Any existing NLP (Natural Language Processing) method can be used here to implement layout analysis to obtain the POI name, which will not be elaborated here.
[0054] S305. Determine the first-level industry category for verification of the POI according to the POI name, where the first-level industry category for verification is used to verify the accuracy of the target first-level industry category.
[0055] Specifically, first, according to the industry category label extraction method in the prior art, the primary industry category of the POI can be determined based on the POI name. Then, the primary industry category is compared with the target primary industry category obtained based on multimodal information. If the two are the same, the accuracy of the target primary industry category can be further determined, thereby achieving the purpose of verifying the accuracy of the target primary industry category.
[0056] It should also be noted that if the primary industry category used for verification is different from the target primary industry category, the target primary industry category can also be used as the final industry category label because the method for determining the industry category based on multimodal information in the embodiments of the present disclosure has a higher confidence level, and at the same time, it can also improve the recall rate of the industry category label.
[0057] S306. Through the layout analysis, obtain the business scope of the POI from the text information.
[0058] S307. In response to the primary industry category used for verification being the same as the target primary industry category, determine the secondary industry category of the POI according to the business scope.
[0059] When the primary industry category used for verification is the same as the target primary industry category, since the target primary industry category has a higher confidence level and accuracy, it indicates that the recognition result of the signboard image also has a certain degree of accuracy. Therefore, the secondary industry category of the POI can be further determined according to the business scope text information therein, thereby improving the industry category of the POI and enhancing its integrity.
[0060] Exemplarily, Figure 3bIt is a schematic diagram of a method for obtaining the industry category of a POI according to an embodiment of the present disclosure. According to the POI signboard image in the figure, the name of the POI is "XXX Studio", and there is a line of small characters "Nail and Eyelash Training Center" in the lower right corner. However, if a semantic analysis is performed based on the POI name according to the prior art, it is impossible to conclude that the first-level industry category of the POI is related to "beauty", let alone that its second-level industry category is "nail art". However, it can be found from the door face image that it contains a female portrait. Therefore, in the embodiment of the present disclosure, the multimodal information of the image and text in the door face image and the signboard image is combined, and the industry category tag Tag related to beauty can be recalled through a cross-modal image and text recall model. At the same time, OCR recognition and layout analysis are performed on the sign image to obtain the POI name and business category information. After an industry category is obtained according to the POI name through the NLP text comprehensive judgment module, on the one hand, the industry category can be used to test the accuracy of the Tag obtained using the cross-modal image and text recall model. On the other hand, the secondary industry category Tag of the POI can be further obtained as "manicure" based on the business scope information. The final industry category Tag is "beauty-manicure", making the industry category information of the POI more complete and accurate.
[0061] The technical solution of the disclosed embodiment combines the multimodal information of the POI's signboard image and storefront image, utilizes a cross-modal image-text recall model to obtain the primary industry category, and then verifies the accuracy of the primary industry category based on the POI name in the signboard image, and further obtains the secondary industry category of the POI through the business scope text information in the signboard image, which can not only ensure a high recall rate of the POI industry category, but also improve the accuracy and completeness of the recalled industry category.
[0062] Figure 4 1 is a schematic diagram of the structure of a device for obtaining the industry category of a POI according to an embodiment of the present disclosure. This embodiment can be applied to the case of obtaining the industry category to which a POI belongs, and relates to the field of image processing technology, especially to intelligent transportation, NLP and deep learning technology. The device can implement the method for obtaining the industry category of a POI described in any embodiment of the present disclosure. Figure 4 As shown, the device 400 specifically includes:
[0063] An image acquisition module 401 is used to acquire a signboard image and a door image of a POI;
[0064] The industry category acquisition module 402 is used to use a pre-trained cross-modal image-text recall model to output the target primary industry category of the POI according to the signboard image and the door image.
[0065] Optionally, the industry category acquisition module 402 includes:
[0066] A scoring unit, configured to score industry categories in an industry category database according to the signboard image and the facade image by using a pre-trained cross-modal image-text retrieval model;
[0067] A candidate industry category set obtaining unit, configured to output a candidate industry category set including a set number of candidate industry categories according to the scoring result, where the candidate industry categories in the candidate industry category set include first-level categories and / or second-level categories;
[0068] A target first-level industry category determining unit, configured to, if there is a same first-level category corresponding to a set proportion of candidate industry categories in the candidate industry category set, use this first-level category as the target first-level industry category of the POI.
[0069] Optionally, the cross-modal image-text retrieval model is a two-tower model, and the two-tower model includes an image tower for obtaining image features in the signboard image and the facade image, and a text tower for obtaining text features in the signboard image and the facade image.
[0070] Optionally, the apparatus further includes a model training module, specifically configured to:
[0071] Obtain POI sample data, where the sample data includes a signboard image, a facade image, and a corresponding industry category of the POI, and the sample data is divided into positive samples and negative samples;
[0072] Input the sample data into a pre-built cross-modal image-text retrieval model, and respectively extract image features and text features in the signboard image and the facade image through the image tower and the text tower in the cross-modal image-text retrieval model, and obtain vector representations of the image features and the text features;
[0073] During the training process, use a preset loss function to make the distance between the vector representations of the image features and the text features and the vector representation of the corresponding industry category in the positive sample tend to approach, and the distance between the vector representations of the image features and the text features and the vector representation of the corresponding industry category in the negative sample tend to be far away.
[0074] Optionally, the image acquisition module 401 is specifically configured to:
[0075] Determine the coordinates of the signboard image on the original image of the POI;
[0076] According to the coordinates, obtain an environmental image within a set range around the signboard image in the original image, and use the environmental image as the facade image of the POI, where the set range is determined according to the length and width of the signboard image.
[0077] Optionally, the apparatus further includes a signboard image processing module, including:
[0078] A text information extraction unit for extracting text information from the signboard image;
[0079] A POI name acquisition unit for obtaining the POI name from the text information by performing layout analysis on the signboard image;
[0080] An inspection unit for determining the primary industry category for inspection of the POI according to the POI name, wherein the primary industry category for inspection is used to inspect the accuracy of the target primary industry category.
[0081] Optionally, the signboard image processing module further includes:
[0082] A business scope acquisition unit for obtaining the business scope of the POI from the text information by means of the layout analysis;
[0083] A secondary industry category acquisition unit for determining the secondary industry category of the POI according to the business scope in response to the primary industry category for inspection being the same as the target primary industry category.
[0084] The above product can execute the method provided in any embodiment of the present disclosure, and has corresponding function modules and beneficial effects for executing the method.
[0085] In the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved are all in compliance with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0086] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0087] Figure 5 The schematic block diagram of an example electronic device 500 that can be used to implement the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0088] As Figure 5As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0089] Multiple components in device 500 are connected to the I / O interface 505, including: an input unit 506, such as a keyboard, a mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, an optical disc, etc.; and a communication unit 509, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0090] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the method for obtaining the industry category of a POI. For example, in some embodiments, the method for obtaining the industry category of a POI can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the method for obtaining the industry category of a POI described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the method for obtaining the industry category of a POI in any other appropriate manner (e.g., by means of firmware).
[0091] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0092] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing device, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can execute entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0093] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0094] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0095] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0096] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services. The server can also be a server of a distributed system, or a server combined with blockchain.
[0097] Artificial intelligence is a discipline that studies to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has both hardware - level technologies and software - level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, machine learning / deep learning technology, big data processing technology, and knowledge graph technology.
[0098] Cloud computing refers to a technical system that accesses an elastic and scalable shared physical or virtual resource pool through a network. The resources can include servers, operating systems, networks, software, applications, and storage devices, etc., and the resources can be deployed and managed in a on-demand and self-service manner. Through cloud computing technology, it can provide efficient and powerful data processing capabilities for the application and model training of technologies such as artificial intelligence and blockchain.
[0099] It should be understood that various forms of processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions provided in this disclosure can be achieved, and no limitations are imposed herein.
[0100] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for obtaining the industry category of a POI, comprising: Obtaining the signboard image and the facade image of the POI; Using a pre-trained cross-modal image-text retrieval model to score the industry categories in the industry category database according to the signboard image and the facade image; According to the result of the scoring, outputting a candidate industry category set including a set number of candidate industry categories, wherein the candidate industry categories in the candidate industry category set include first-level categories and / or second-level categories; If, in the candidate industry category set, there is a set proportion of candidate industry categories corresponding to the same first-level category, then taking this first-level category as the target first-level industry category of the POI; Wherein, the cross-modal image-text retrieval model is a two-tower model, and the two-tower model includes an image tower for obtaining image features in the signboard image and the facade image, and a text tower for obtaining text features in the signboard image and the facade image; The training process of the cross-modal image-text retrieval model includes: Obtaining POI sample data, wherein the sample data includes the signboard image, the facade image and the corresponding industry category of the POI, and the sample data is divided into positive samples and negative samples; Inputting the sample data into a pre-built cross-modal image-text retrieval model, and respectively extracting the image features and text features in the signboard image and the facade image through the image tower and the text tower in the cross-modal image-text retrieval model, and obtaining the vector representations of the image features and the text features; During the training process, using a preset loss function to make the distance between the vector representations of the image features and the text features and the vector representation of the corresponding industry category in the positive sample tend to approach, and the distance from the vector representation of the corresponding industry category in the negative sample tend to be far away.
2. The method according to claim 1, wherein, The obtaining of the facade image of the POI includes: Determining the coordinates of the signboard image on the original image of the POI; According to the coordinates, obtaining the environmental image within a set range around the signboard image in the original image, and taking the environmental image as the facade image of the POI, wherein the set range is determined according to the length and width of the signboard image.
3. The method according to claim 1, further comprising: Extracting the text information in the signboard image; Through layout analysis of the signboard image, obtaining the POI name from the text information; Determining the first-level industry category for verification of the POI according to the POI name, wherein the first-level industry category for verification is used to verify the accuracy of the target first-level industry category.
4. The method according to claim 3, further comprising: Through the layout analysis, obtaining the business scope of the POI from the text information; In response to the first-level industry category for verification being the same as the target first-level industry category, determining the second-level industry category of the POI according to the business scope.
5. An apparatus for obtaining the industry category of a POI, comprising: An image acquisition module for obtaining the signboard image and the facade image of the POI; An industry category acquisition module, configured to use a pre-trained cross-modal image-text retrieval model to output a target primary industry category of the POI according to the signboard image and the facade image; Wherein, the industry category acquisition module includes: A scoring unit, configured to use a pre-trained cross-modal image-text retrieval model to score the industry categories in the industry category database according to the signboard image and the facade image; A candidate industry category set acquisition unit, configured to output a candidate industry category set including a set number of candidate industry categories according to the scoring result, wherein the candidate industry categories in the candidate industry category set include primary categories and / or secondary categories; A target primary industry category determination unit, configured to, if there is a same primary category corresponding to a set proportion of the candidate industry categories in the candidate industry category set, use this primary category as the target primary industry category of the POI; The cross-modal image-text retrieval model is a two-tower model, and the two-tower model includes an image tower for obtaining image features in the signboard image and the facade image, and a text tower for obtaining text features in the signboard image and the facade image; The device further includes a model training module, specifically configured to: Obtain POI sample data, where the sample data includes a signboard image, a facade image, and a corresponding industry category of the POI, and the sample data is divided into positive samples and negative samples; Input the sample data into a pre-built cross-modal image-text retrieval model, and respectively extract image features and text features in the signboard image and the facade image through the image tower and the text tower in the cross-modal image-text retrieval model, and obtain vector representations of the image features and the text features; During the training process, use a preset loss function to make the distance between the vector representations of the image features and the text features and the vector representation of the corresponding industry category in the positive sample tend to approach, and the distance from the vector representation of the corresponding industry category in the negative sample tend to be far away.
6. An electronic device, including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method for acquiring the industry category of the POI according to any one of claims 1-4.
7. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method for acquiring the industry category of the POI according to any one of claims 1-4.
8. A computer program product, including a computer program, where the computer program, when executed by a processor, implements the method for acquiring the industry category of the POI according to any one of claims 1-4.
Citation Information
Patent Citations
Image category identification method and device
CN110705460A
Interest point information processing method and device, electronic equipment and storage medium
CN111832578A