Indoor area search method, device, electronic device and storage medium
By acquiring a set of candidate images and using a large model to identify the target image, combined with cloud services and smart terminals, automatic indoor area positioning is achieved, solving the problems of low efficiency and insufficient accuracy in existing technologies, and providing an efficient and accurate indoor area retrieval solution.
Patent Information
- Application Number
- CN202411545905.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2044-10-31
AI Technical Summary
Existing indoor area retrieval methods are inefficient and difficult to guarantee accuracy, and are unable to locate indoor areas through pre-marked geographic coordinates.
By obtaining a set of candidate images and using a large model to identify target images that match the retrieval requirement description information, the retrieval results for indoor areas are generated. In combination with the collaborative work of cloud services and smart terminals, indoor area positioning is automatically processed.
It simplifies manual operations, improves the processing efficiency and accuracy of indoor area retrieval, and provides flexible input methods and accurate location information display.
Smart Images

Figure CN119474443B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to indoor area retrieval methods, devices, electronic devices, and storage media in the fields of large models, deep learning, computer vision, and cloud computing. Background Art
[0002] Outdoor area searches can often be performed using pre-marked geographic coordinates. For example, to find a specific park, one can simply enter the park name on an electronic map. However, indoor areas are often more complex and cannot be pre-marked with geographic coordinates. Therefore, current indoor area searches are mainly performed through human interaction, which is inefficient and difficult to guarantee accuracy. Summary of the Invention
[0003] The present disclosure provides an indoor area search method, device, electronic device, and storage medium.
[0004] An indoor area retrieval method, comprising:
[0005] Acquire retrieval requirement description information, where the retrieval requirement description information is description information about a target area to be retrieved, and the target area is an indoor area;
[0006] A first image set is formed using the acquired candidate images, where the number of the candidate images is greater than one, and the candidate images are images captured covering an indoor area of a target building, where the target building is a building to be searched for in this search;
[0007] In response to identifying a candidate image matching the retrieval requirement description information from the first image set, the matching candidate image is used as a target image, and a retrieval result corresponding to the target area is generated and returned based on the target image.
[0008] An indoor area retrieval method, comprising:
[0009] Acquire retrieval requirement description information input by a user, wherein the retrieval requirement description information is description information about a target area to be retrieved, and the target area is an indoor area;
[0010] Sending the search requirement description information to the cloud service;
[0011] In response to obtaining the retrieval result returned by the cloud service, the retrieval result is displayed to the user, or the retrieval result is displayed to the user after a predetermined conversion, the retrieval result is a retrieval result corresponding to the target area generated by the cloud service based on the target image, the target image is a candidate image identified by the cloud service from the first image set that matches the retrieval requirement description information, the first image set is composed of candidate images, and the candidate images in the first image set are images captured covering the indoor area of the target building, and the target building is the building that serves as the retrieval object of this retrieval.
[0012] An indoor area search device includes: a first acquisition module, a second acquisition module and a search processing module;
[0013] The first acquisition module is used to acquire retrieval requirement description information, wherein the retrieval requirement description information is description information about a target area to be searched, and the target area is an indoor area;
[0014] The second acquisition module is configured to form a first image set using the acquired candidate images, where the number of the candidate images is greater than one, and the candidate images are images captured covering an indoor area of a target building, where the target building is a building to be searched for in this search;
[0015] The retrieval processing module is used to respond to identifying a candidate image matching the retrieval requirement description information from the first image set, taking the matching candidate image as a target image, and generating a retrieval result corresponding to the target area based on the target image and returning it.
[0016] An indoor area search device includes: a third acquisition module, an information sending module, and a result processing module;
[0017] The third acquisition module is configured to acquire retrieval requirement description information input by a user, wherein the retrieval requirement description information is description information about a target area to be searched, and the target area is an indoor area;
[0018] The information sending module is used to send the search requirement description information to the cloud service;
[0019] The result processing module is used to display the retrieval result to the user in response to obtaining the retrieval result returned by the cloud service, or to display the retrieval result to the user after performing a predetermined conversion on the retrieval result, wherein the retrieval result is a retrieval result corresponding to the target area generated by the cloud service based on the target image, and the target image is a candidate image that matches the retrieval requirement description information identified by the cloud service from the first image set, the first image set is composed of candidate images, and the candidate images in the first image set are images captured covering the indoor area of the target building, and the target building is the building that serves as the retrieval object of this retrieval.
[0020] An electronic device, comprising:
[0021] at least one processor; and
[0022] a memory communicatively connected to the at least one processor; wherein,
[0023] The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the method as described above.
[0024] A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0025] A computer program product comprises a computer program / instruction, which implements the above method when executed by a processor.
[0026] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0028] Figure 1 This is a flowchart of the first embodiment of the indoor area search method disclosed in the present invention;
[0029] Figure 2 This is a flowchart of the second embodiment of the indoor area retrieval method described in the present disclosure;
[0030] Figure 3 Schematic diagram of the structure of the first embodiment 300 of the indoor area search device disclosed in the present invention;
[0031] Figure 4Schematic diagram of the structure of the second embodiment 400 of the indoor area search device described in the present disclosure;
[0032] Figure 5 FIG. 5 is a schematic block diagram of an electronic device 500 that can be used to implement an embodiment of the present disclosure. DETAILED DESCRIPTION
[0033] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0034] Furthermore, it should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " as used herein generally indicates that the associated objects are in an "or" relationship.
[0035] Figure 1 This is a flow chart of the first embodiment of the indoor area search method described in this disclosure. Figure 1 As shown, the following specific implementation methods are included.
[0036] In step 101, retrieval requirement description information is obtained. The retrieval requirement description information is description information about a target area to be retrieved, and the target area is an indoor area.
[0037] In step 102, the acquired candidate images are used to form a first image set. The number of candidate images is greater than 1, and the candidate images are images captured covering the indoor area of the target building. The target building is the building that is the retrieval object of this retrieval.
[0038] In step 103, in response to identifying a candidate image matching the retrieval requirement description information from the first image set, the matching candidate image is used as a target image, and a retrieval result corresponding to the target area is generated and returned based on the target image.
[0039] By adopting the scheme described in the above method embodiment, the retrieval results (i.e., search results) of the target area to be retrieved (i.e., to be searched) can be automatically generated based on the retrieval requirement description information and candidate images through matching and recognition operations. The target area is an indoor area, thereby simplifying manual operations, improving processing efficiency, and improving the accuracy of processing results.
[0040] In practical applications, Figure 1 The execution subject of the illustrated embodiment may be a cloud service, ie, a background service. The cloud service may obtain retrieval requirement description information from the smart terminal and may obtain candidate images captured by the camera.
[0041] For example, if a user is in a shopping mall (target building) and wants to find a specific restaurant, or generally wants to find restaurants in the mall, they can enter a search requirement description into the smart terminal, such as "Help me find ** restaurant." Another example is if a user is in a supermarket (target building) and wants to find the fresh food section, they can enter a search requirement description into the smart terminal, such as "Help me find the location of the fresh food section."
[0042] In addition, at least one camera can be installed on each floor of the target building. The specific number can be determined according to actual needs. Each camera can collect candidate images. Generally speaking, the candidate images collected by the cameras on each floor need to be able to cover all indoor areas of the target building.
[0043] In some embodiments of the present disclosure, the candidate images constituting the first image set may be the most recently acquired candidate images, wherein the candidate images may be acquired again after every predetermined time period.
[0044] In other words, each camera can recapture images at predetermined intervals and send them to the backend service. The specific value of the predetermined interval can also be determined based on actual needs. For example, taking a shopping mall as the target building, considering that the store layout and other factors within the mall may change, candidate images can be collected periodically. The most recently acquired candidate images can be used to form the first image set to improve the accuracy of subsequent processing results.
[0045] For the first set of images, candidate images matching the search requirement description information can be identified and used as target images. In actual applications, it is possible that no target images matching the search requirement description information can be identified. For example, a user may want to search for a certain brand of fast food restaurant, but no such restaurant exists in the target building. In this case, there are no restrictions on how to handle it. For example, a prompt message stating "No relevant results found" can be directly returned to the smart terminal for display to the user.
[0046] In some embodiments of the present disclosure, in response to determining that the retrieval requirement description information is in text form, keywords can be extracted from the retrieval requirement description, and the target image can be identified based on the keywords. In response to determining that the retrieval requirement description information is in voice form, the retrieval requirement description information can be converted from voice to text, and keywords can be extracted from the conversion result, and then the target image can be identified based on the keywords.
[0047] In other words, it can support users to input the description information of search requirements in text form, and it can also support users to input the description information of search requirements in voice form. It is very flexible and convenient, and can meet the different usage needs of users.
[0048] If it is in text form, keywords can be directly extracted from the retrieval requirement description information. If it is in voice form, the retrieval requirement description information can first be converted from voice to text to obtain the required conversion result, and then keywords can be extracted from the conversion result.
[0049] The extracted keywords may include the name and characteristics of the target area. For example, if the search requirement description is "Help me find the location of the fresh food section," the extracted keyword may be "fresh food section" (the name of the target area). For another example, if the search requirement description is "Help me find a hot pot restaurant in this mall," the extracted keywords may be "hot pot" and "restaurant."
[0050] In addition, there is no restriction on how to extract keywords. For example, the extraction can be performed based on rules, or the extraction can be performed using a pre-trained extraction model.
[0051] After obtaining the keywords, the target image can be identified based on the keywords. In some embodiments of the present disclosure, the following processing can be performed for each candidate image: in response to identifying that the candidate image includes the target area based on the keywords, the candidate image is added to the second image set; in response to determining that the second image set is not empty, the candidate image in the second image set is determined to be the target image.
[0052] For example, assuming that the first image set includes 30 candidate images, for ease of description, they are referred to as candidate images 1 to candidate images 30. For candidate image 1, if it is identified based on the keyword that candidate image 1 includes the target area, then candidate image 1 can be added to the second image set. If it is identified based on the keyword that candidate image 1 does not include the target area, then candidate image 1 will not be added to the second image set. Similarly, for candidate image 2, if it is identified based on the keyword that candidate image 2 includes the target area, then candidate image 2 can also be added to the second image set. If it is identified based on the keyword that candidate image 2 does not include the target area, then candidate image 2 will not be added to the second image set. The processing method for other candidate images is similar and will not be repeated. Assuming that the second image set is not empty after all candidate images have been processed, the candidate images in the second image set can be determined as target images.
[0053] By adopting the above processing method, all target images can be identified from the first image set, thereby improving the comprehensiveness and accuracy of the recognition result.
[0054] Alternatively, in some embodiments of the present disclosure, when performing target image recognition based on keywords, a statistical parameter may be set, with an initial value of 0, and each candidate image may be traversed in sequence according to a predetermined order, and the following processing may be performed for each candidate image traversed: in response to recognizing that the candidate image includes a target area based on the keyword, the candidate image is added to the third image set, and the value of the statistical parameter is increased by 1; in response to determining that the latest value of the statistical parameter is equal to N, the traversal is ended, and the candidate image in the third image set is determined as the target image, where N is a positive integer; in response to determining that the latest value of the statistical parameter is equal to The value is less than N, and it is determined that there are candidate images that have not been traversed, continuing to traverse the next candidate image, in response to determining that the latest value of the statistical parameter is less than N, and determining that all candidate images have been traversed, the candidate image in the third image set is determined as the target image; in response to identifying that the candidate image does not include the target area according to the keyword, and determining that there are candidate images that have not been traversed, continuing to traverse the next candidate image, in response to identifying that the candidate image does not include the target area according to the keyword, and determining that all candidate images have been traversed, the candidate image in the third image set is determined as the target image.
[0055] The specific value of N can be determined according to actual needs, and N can represent an upper limit set by the user. In addition, the specific order of the predetermined order can also be determined according to actual needs.
[0056] Assume that the first image set includes a total of 30 candidate images, which are respectively referred to as candidate image 1 to candidate image 30 for ease of expression, and assume that they are traversed in the order of candidate image 1 to candidate image 30. Taking candidate image 15 as an example, if the target area is included in candidate image 15 according to the keyword, then candidate image 15 can be added to the third image set, and the value of the statistical parameter can be increased by 1. If it is determined that the value of the updated statistical parameter is equal to N, candidate images 16 to candidate image 30 can be not traversed, but the candidate image in the third image set can be directly determined as the target image. If it is determined that the value of the updated statistical parameter is less than N, then candidate image 16 can be continued to be traversed. Alternatively, if the target area is not included in candidate image 15 according to the keyword, candidate image 16 can be continued to be traversed. If after all 30 candidate images are traversed, the value of the statistical parameter is less than or equal to N, then the candidate image in the third image set can also be determined as the target image.
[0057] Using the above processing method, once it is determined that the number of identified candidate images including the target area reaches N, the traversal can be ended and the candidate image in the current third image set can be directly determined as the target image, so there is no need to identify all candidate images, thereby improving processing efficiency and avoiding returning too much information, affecting the user's information browsing efficiency, etc.
[0058] For any candidate image, the large model can be used to identify whether it includes the target area. In some embodiments of the present disclosure, a prompt for the large model can be first constructed based on the keyword. The candidate image and the prompt can then be used as input to the large model. In response to obtaining a first output result of the large model, it can be determined that the candidate image includes the target area. In response to obtaining a second output result of the large model, it can be determined that the candidate image does not include the target area.
[0059] There are no restrictions on how to construct the prompt word. For example, if the extracted keywords include "hot pot" and "restaurant," the constructed prompt word could be "You are ***, please determine whether the image contains a hot pot restaurant, ****." Accordingly, the large model can generate a first output result indicating "yes" or a second output result indicating "no" based on the input candidate image and prompt word.
[0060] The big model refers to the Large Language Model (LLM), which is a deep learning model trained using a large amount of text data. It can generate natural language text or understand the meaning of language text, and is an important path to artificial intelligence.
[0061] Considering that large models can understand images and, by extension, objects in geographic space, the solution described in this disclosure utilizes large models to identify whether candidate images contain target areas, thereby improving the accuracy of recognition results. Furthermore, the solution described in this disclosure imposes no restrictions on the type of large model, as long as it can perform the corresponding function.
[0062] After the target image is determined, a search result corresponding to the target area can be generated based on the target image, and the search result can be returned to the smart terminal.
[0063] In some embodiments of the present disclosure, the first image set may also include: a camera identifier (ID) of a camera corresponding to each candidate image, and the camera corresponding to any candidate image is the camera used to capture the candidate image. Accordingly, a method of returning a retrieval result corresponding to a target area generated according to a target image may include: determining the location information corresponding to the target image as the target location information by querying the correspondence between each pre-saved camera identifier and the corresponding location information, and returning the target location information as a retrieval result, wherein any location information includes: the floor where the corresponding camera is located and the coordinate information of the corresponding camera on the floor where it is located.
[0064] Each camera sends its camera ID when sending candidate images to the backend service. Furthermore, the cloud service pre-stores location information corresponding to different camera IDs. For each camera, the corresponding location information may include the floor on which the camera is located, such as the 2nd floor, and the camera's coordinates within that floor. These coordinates may refer to specific horizontal and vertical coordinates within the building.
[0065] For example, assuming that the target image only includes candidate image 15, the location information corresponding to the camera corresponding to candidate image 15 can be determined by querying the corresponding relationship, and the location information can be returned to the smart terminal as a retrieval result, so that the smart terminal can display the retrieval result or the retrieval result after a predetermined conversion to the user, so that the user can know the location of the target area, etc.
[0066] In addition, in actual applications, cloud services can be further subdivided into access services, voice conversion services, extraction services, retrieval services, and database services (Mongo) based on distributed file storage. The access service can obtain retrieval requirement description information from smart terminals and candidate images from cameras, and can return retrieval results to smart terminals. Other services can also be called, such as calling the voice conversion service to convert the voice-based retrieval requirement description information into text, calling the extraction service to extract keywords from the text-based retrieval requirement description information, calling the retrieval service to identify the target image from the first image set based on the keywords and return the retrieval results generated based on the target image, and calling the Mongo service to query the correspondence between each camera identifier and the corresponding location information. The access service, voice conversion service, extraction service, and retrieval service can all be implemented using serverless technology.
[0067] Furthermore, the search requirement description information obtained by the cloud service may not come from the user, but from smart devices such as delivery robots. In this case, the search requirement description information is usually in text form and can be sent to the cloud service through Hypertext Transfer Protocol (HTTP) instead of the smart terminal. After generating the search results, the cloud service can return the search results to the smart device.
[0068] The above mainly describes the solution of the present disclosure from the perspective of the cloud service side. The following further describes the solution of the present disclosure from the perspective of the smart terminal side.
[0069] Figure 2 This is a flow chart of the second embodiment of the indoor area search method described in the present disclosure. Figure 2 As shown, the following specific implementation methods are included.
[0070] In step 201, the retrieval requirement description information input by the user is obtained. The retrieval requirement description information is description information about a target area to be retrieved, and the target area is an indoor area.
[0071] In step 202, the search requirement description information is sent to the cloud service.
[0072] In step 203, in response to obtaining the retrieval result returned by the cloud service, the retrieval result is displayed to the user, or the retrieval result is displayed to the user after a predetermined conversion. The retrieval result is a retrieval result corresponding to the target area generated by the cloud service based on the target image. The target image is a candidate image that matches the retrieval requirement description information identified by the cloud service from the first image set. The first image set is composed of candidate images, and the candidate images in the first image set are images captured covering the indoor area of the target building. The target building is the building that is the retrieval object of this retrieval.
[0073] By adopting the scheme described in the above method embodiment, the retrieval results of the target area to be retrieved can be automatically generated based on the retrieval requirement description information and candidate images through matching and recognition operations. The target area is an indoor area, thereby simplifying manual operations, improving processing efficiency, and improving the accuracy of processing results.
[0074] In practical applications, Figure 2 The execution subject of the embodiment shown may be a smart terminal, which may obtain the search requirement description information input by the user and send it to the cloud service.
[0075] The smart terminal may also obtain search results generated and returned by the cloud service. In response to identifying a candidate image from the first image set that matches the search requirement description information, the cloud service may use the matching candidate image as a target image and generate and return a search result corresponding to the target area based on the target image.
[0076] In some embodiments of the present disclosure, the first image set may also include: camera identifiers of cameras corresponding to each candidate image, the cameras corresponding to any candidate image are cameras used to capture the candidate image, and at least one camera is provided on each floor of the target building. Accordingly, the retrieval results may include: the cloud service determines the location information corresponding to the target image by querying the correspondence between each pre-saved camera identifier and the corresponding location information, wherein any location information includes: the floor where the corresponding camera is located and the coordinate information of the corresponding camera on the floor.
[0077] Furthermore, the intelligent terminal can directly display the search results to the user, or can display the search results to the user after performing a predetermined conversion. For example, assuming that the target image only includes candidate image 15, the location information corresponding to the camera corresponding to candidate image 15 can be directly displayed to the user. Alternatively, the location information corresponding to the camera corresponding to candidate image 15 can be first converted into a form that is easier for the user to understand, and then the conversion result can be displayed to the user. For example, if the location information corresponding to the camera corresponding to candidate image 15 is the second floor + coordinate information, the coordinate information can be converted into a form that is easier for the user to understand, such as the southwest corner, the middle of the north side, etc., to facilitate user use.
[0078] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should know that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure. In addition, for parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0079] The above is an introduction to the method embodiment. The following is a further explanation of the solution disclosed in the present disclosure through an apparatus embodiment.
[0080] Figure 3 FIG. 3 is a schematic diagram of the structure of the first embodiment 300 of the indoor area search device described in the present disclosure. Figure 3 As shown, it includes: a first acquisition module 301, a second acquisition module 302 and a retrieval processing module 303.
[0081] The first acquisition module 301 is used to acquire retrieval requirement description information, where the retrieval requirement description information is description information about a target area to be retrieved, and the target area is an indoor area.
[0082] The second acquisition module 302 is used to use the acquired candidate images to form a first image set, where the number of candidate images is greater than 1, and the candidate images are images captured covering the indoor area of the target building, which is the building serving as the retrieval object of this retrieval.
[0083] The retrieval processing module 303 is configured to, in response to identifying a candidate image matching the retrieval requirement description information from the first image set, use the matching candidate image as a target image, and generate a retrieval result corresponding to the target area based on the target image and return it.
[0084] In some embodiments of the present disclosure, the candidate images constituting the first image set may be the most recently acquired candidate images, wherein the second acquisition module 302 may reacquire the candidate images every predetermined period of time.
[0085] For the first image set, the retrieval processing module 303 may identify candidate images that match the retrieval requirement description information and use the identified matching candidate images as target images. In some embodiments of the present disclosure, in response to determining that the retrieval requirement description information is in text form, the retrieval processing module 303 may extract keywords from the retrieval requirement description and identify the target image based on the keywords. In response to determining that the retrieval requirement description information is in voice form, the retrieval processing module 303 may convert the retrieval requirement description information from voice to text, extract keywords from the conversion result, and then identify the target image based on the keywords.
[0086] In some embodiments of the present disclosure, the retrieval processing module 303 may perform the following processing for each candidate image: in response to identifying that the candidate image includes a target area based on keywords, the candidate image is added to the second image set; in response to determining that the second image set is not empty, the candidate image in the second image set is determined as the target image.
[0087] Alternatively, in some embodiments of the present disclosure, the retrieval processing module 303 may set a statistical parameter with an initial value of 0, and may traverse each candidate image in a predetermined order, and perform the following processing on each candidate image traversed: in response to identifying that the candidate image includes a target area based on the keyword, the candidate image is added to the third image set, and the value of the statistical parameter is increased by 1; in response to determining that the latest value of the statistical parameter is equal to N, the traversal is ended, and the candidate image in the third image set is determined as the target image, N is a positive integer; in response to determining that the latest value of the statistical parameter is less than N, and it is determined that there are candidate images that have not been traversed, continue to traverse the next candidate image, in response to determining that the latest value of the statistical parameter is less than N, and it is determined that all candidate images have been traversed, determine the candidate image in the third image set as the target image; in response to identifying that the candidate image does not include the target area according to the keyword, and it is determined that there are candidate images that have not been traversed, continue to traverse the next candidate image, in response to identifying that the candidate image does not include the target area according to the keyword, and it is determined that all candidate images have been traversed, determine the candidate image in the third image set as the target image.
[0088] For any candidate image, the retrieval processing module 303 can use the large model to identify whether it includes the target area. In some embodiments of the present disclosure, the retrieval processing module 303 can first construct a prompt word for the large model based on the keyword, and then use the candidate image and the prompt word as input to the large model. In response to obtaining a first output result of the large model, it can be determined that the candidate image includes the target area. In response to obtaining a second output result of the large model, it can be determined that the candidate image does not include the target area.
[0089] In some embodiments of the present disclosure, the first image set may also include: camera identifiers of cameras corresponding to each candidate image, and the cameras corresponding to any candidate image are cameras used to capture the candidate image. Accordingly, the retrieval processing module 303 generates retrieval results corresponding to the target area based on the target image and returns the results in a manner that may include: determining the location information corresponding to the target image as the target location information by querying the correspondence between each pre-saved camera identifier and the corresponding location information, and returning the target location information as the retrieval result, wherein any location information includes: the floor where the corresponding camera is located and the coordinate information of the corresponding camera on the floor where it is located.
[0090] Figure 4 FIG. 4 is a schematic diagram of the structure of the second embodiment 400 of the indoor area search device described in the present disclosure. Figure 4 As shown, it includes: a third acquisition module 401, an information sending module 402 and a result processing module 403.
[0091] The third acquisition module 401 is used to acquire the search requirement description information input by the user. The search requirement description information is description information about the target area to be searched, and the target area is an indoor area.
[0092] The information sending module 402 is used to send the search requirement description information to the cloud service.
[0093] The result processing module 403 is used to display the search results to the user in response to obtaining the search results returned by the cloud service, or to display the search results to the user after performing a predetermined conversion on the search results. The search results are search results corresponding to the target area generated by the cloud service based on the target image. The target image is a candidate image that matches the search requirement description information identified by the cloud service from the first image set. The first image set is composed of candidate images, and the candidate images in the first image set are images captured and covering the indoor area of the target building. The target building is the building that is the search object of this search.
[0094] In some embodiments of the present disclosure, the first image set may also include: camera identifiers of cameras corresponding to each candidate image, the cameras corresponding to any candidate image are cameras used to capture the candidate image, and at least one camera is provided on each floor of the target building. Accordingly, the retrieval results may include: the cloud service determines the location information corresponding to the target image by querying the correspondence between each pre-saved camera identifier and the corresponding location information, wherein any location information includes: the floor where the corresponding camera is located and the coordinate information of the corresponding camera on the floor.
[0095] The specific working processes of the above-mentioned device embodiments can refer to the relevant descriptions in the above-mentioned method embodiments and will not be repeated here.
[0096] The solutions described in this disclosure can be applied to the field of artificial intelligence, particularly in the areas of large models, deep learning, computer vision, and cloud computing. Artificial intelligence is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It involves both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technology.
[0097] Furthermore, the search requirement description information and images in the embodiments of this disclosure are not targeted at any specific user and do not reflect any specific user's personal information. The collection, storage, use, processing, transmission, provision, and disclosure of user personal information in the technical solutions of this disclosure comply with relevant laws and regulations and do not violate public order and good morals.
[0098] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0099] Figure 5 A schematic block diagram of an electronic device 500 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0100] like Figure 5 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0101] Various components in device 500 are connected to I / O interface 505, including: an input unit 506, such as a keyboard, mouse, etc.; an output unit 507, such as various types of displays, speakers, etc.; a storage unit 508, such as a magnetic disk, optical disk, etc.; and a communication unit 509, such as a network card, modem, wireless communication transceiver, etc. The communication unit 509 allows device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0102] The computing unit 501 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI, Artificial Intelligence) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSP, Digital Signal Processing), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 501 performs the various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the methods described in the present disclosure can be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to execute the method described in the present disclosure in any other appropriate manner (for example, by means of firmware).
[0103] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0104] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0105] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0106] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0107] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0108] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0109] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0110] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for indoor area retrieval, comprising: Acquire retrieval requirement description information, where the retrieval requirement description information is description information about a target area to be retrieved, and the target area is an indoor area; A first image set is formed using the acquired candidate images, where the number of the candidate images is greater than one, and the candidate images are images captured covering an indoor area of a target building, where the target building is a building to be searched for in this search; In response to identifying a candidate image matching the search requirement description information from the first image set, taking the matching candidate image as a target image, including: in response to determining that the search requirement description information is in text form, extracting keywords from the search requirement description, and identifying the target image based on the keywords, the target image being a candidate image including the target area, and generating a search result corresponding to the target area based on the target image and returning it; Among them, the method for determining whether any candidate image includes the target area includes: constructing a prompt word of a large model based on the keyword, taking the candidate image and the prompt word as input of the large model, and in response to obtaining a first output result of the large model, determining that the candidate image includes the target area; and in response to obtaining a second output result of the large model, determining that the candidate image does not include the target area.
2. The method according to claim 1, wherein The candidate images include: the candidate images acquired most recently; wherein the candidate images are acquired again every time a predetermined time period elapses.
3. The method according to claim 1, further comprising: In response to determining that the retrieval requirement description information is in voice form, the retrieval requirement description information is converted from voice to text, and keywords are extracted from the conversion result.
4. The method according to claim 3, wherein: The identifying of the target image according to the keyword includes: For each candidate image, the following processing is performed: in response to identifying that the candidate image includes the target area according to the keyword, the candidate image is added to the second image set; In response to determining that the second image set is not empty, a candidate image in the second image set is determined as the target image.
5. The method according to claim 3, wherein: The identifying of the target image according to the keyword includes: Set the statistical parameters, with the initial value being 0, and traverse each candidate image in the predetermined order. For each candidate image traversed, perform the following processing: In response to identifying that the candidate image includes the target area according to the keyword, adding the candidate image to the third image set and increasing the value of the statistical parameter by 1; in response to determining that the latest value of the statistical parameter is equal to N, ending the traversal and determining the candidate image in the third image set as the target image, where N is a positive integer; in response to determining that the latest value of the statistical parameter is less than N and that there is a candidate image that has not been traversed, continuing to traverse the next candidate image; in response to determining that the latest value of the statistical parameter is less than N and that all candidate images have been traversed, determining the candidate image in the third image set as the target image; In response to identifying that the target area is not included in the candidate image based on the keyword and determining that there is a candidate image that has not been traversed, continuing to traverse the next candidate image; in response to identifying that the target area is not included in the candidate image based on the keyword and determining that all candidate images have been traversed, determining the candidate image in the third image set as the target image.
6. The method according to any one of claims 1 to 5, wherein The first image set also includes: camera identifiers of cameras corresponding to the candidate images, wherein the cameras corresponding to any candidate image are cameras used to capture the candidate images, and each floor of the target building is provided with at least one camera; The method of generating the retrieval result corresponding to the target area based on the target image includes: determining the position information corresponding to the target image as the target position information by querying the correspondence between each pre-saved camera identifier and the corresponding position information, and returning the target position information as the retrieval result, wherein any position information includes: the floor where the corresponding camera is located and the coordinate information of the corresponding camera on the floor where it is located.
7. A method for indoor area retrieval, comprising: Acquire retrieval requirement description information input by a user, wherein the retrieval requirement description information is description information about a target area to be retrieved, and the target area is an indoor area; Sending the search requirement description information to the cloud service; In response to obtaining a search result returned by the cloud service, the search result is displayed to the user, or the search result is displayed to the user after undergoing a predetermined conversion. The search result is a search result corresponding to the target area generated by the cloud service based on the target image. The target image is a candidate image identified by the cloud service from a first image set that matches the search requirement description information. The first image set is composed of candidate images, and the candidate images in the first image set are images captured covering an indoor area of a target building. The target building is a building that is the search object of this search. The target image is a candidate image including the target area identified based on the keywords in response to determining that the search requirement description information is in text form and performing keyword extraction on the search requirement description. For any candidate image, the keywords are used to construct prompt words for a large model, and the candidate image and the prompt words are used as inputs of the large model to obtain a first output result or a second output result of the large model. The first output result is used to indicate that the candidate image includes the target area, and the second output result is used to indicate that the candidate image does not include the target area.
8. The method according to claim 7, wherein: The first image set also includes: camera identifiers of cameras corresponding to the candidate images, wherein the cameras corresponding to any candidate image are cameras used to capture the candidate images, and each floor of the target building is provided with at least one camera; The retrieval results include: the cloud service determines the location information corresponding to the target image by querying the correspondence between each pre-saved camera identification and the corresponding location information, wherein any location information includes: the floor where the corresponding camera is located and the coordinate information of the corresponding camera on the floor.
9. An indoor area search device, comprising: a first acquisition module, a second acquisition module, and a retrieval processing module; The first acquisition module is used to acquire retrieval requirement description information, wherein the retrieval requirement description information is description information about a target area to be searched, and the target area is an indoor area; The second acquisition module is configured to form a first image set using the acquired candidate images, where the number of the candidate images is greater than one, and the candidate images are images captured covering an indoor area of a target building, where the target building is a building to be searched for in this search; The retrieval processing module is used to, in response to identifying a candidate image matching the retrieval requirement description information from the first image set, use the matching candidate image as a target image, including: in response to determining that the retrieval requirement description information is in text form, extracting keywords from the retrieval requirement description, and identifying the target image based on the keywords, the target image being a candidate image including the target area, and generating a retrieval result corresponding to the target area based on the target image and returning it, wherein a method for determining whether the target area is included in any candidate image includes: constructing a prompt word of a large model based on the keywords, using the candidate image and the prompt word as input of the large model, determining that the candidate image includes the target area in response to obtaining a first output result of the large model, and determining that the candidate image does not include the target area in response to obtaining a second output result of the large model.
10. The device according to claim 9, wherein The candidate images include: the most recently acquired candidate image; The second acquisition module is further configured to reacquire the candidate image every time a predetermined time period has passed.
11. The device according to claim 9, wherein The retrieval processing module is further configured to, in response to determining that the retrieval requirement description information is in voice form, convert the retrieval requirement description information from voice to text, and perform keyword extraction on the conversion result.
12. The device according to claim 11, wherein The retrieval processing module performs the following processing on each candidate image: in response to identifying that the candidate image includes the target area based on the keyword, the candidate image is added to the second image set; in response to determining that the second image set is not empty, the candidate image in the second image set is determined as the target image.
13. The device according to claim 11, wherein The retrieval processing module sets a statistical parameter with an initial value of 0, and traverses each candidate image in a predetermined order, and performs the following processing on each traversed candidate image: in response to identifying that the candidate image includes the target area according to the keyword, the candidate image is added to the third image set, and the value of the statistical parameter is increased by 1; in response to determining that the latest value of the statistical parameter is equal to N, the traversal is terminated, and the candidate image in the third image set is determined as the target image, where N is a positive integer; in response to determining that the latest value of the statistical parameter is less than N and that there is a candidate image that has not been traversed, the traversal continues to the next candidate image; in response to determining that the latest value of the statistical parameter is less than N and that all candidate images have been traversed, the candidate image in the third image set is determined as the target image; In response to identifying that the target area is not included in the candidate image based on the keyword and determining that there are candidate images that have not been traversed, continue to traverse the next candidate image. In response to identifying that the target area is not included in the candidate image based on the keyword and determining that all candidate images have been traversed, determine the candidate image in the third image set as the target image.
14. The device according to any one of claims 9 to 13, wherein: The first image set also includes: camera identifiers of cameras corresponding to the candidate images, wherein the cameras corresponding to any candidate image are cameras used to capture the candidate images, and each floor of the target building is provided with at least one camera; The retrieval processing module determines the location information corresponding to the target image as the target location information by querying the correspondence between each pre-saved camera identification and the corresponding location information, and returns the target location information as the retrieval result, wherein any location information includes: the floor where the corresponding camera is located and the coordinate information of the corresponding camera on the floor where it is located.
15. An indoor area search device, comprising: A third acquisition module, an information sending module, and a result processing module; The third acquisition module is configured to acquire retrieval requirement description information input by a user, wherein the retrieval requirement description information is description information about a target area to be searched, and the target area is an indoor area; The information sending module is used to send the search requirement description information to the cloud service; The result processing module is configured to, in response to obtaining a search result returned by the cloud service, display the search result to the user, or display the search result to the user after performing a predetermined conversion on the search result. The search result is a search result corresponding to the target area generated by the cloud service based on the target image. The target image is a candidate image identified by the cloud service from a first image set that matches the search requirement description information. The first image set is composed of candidate images, and the candidate images in the first image set are images captured covering an indoor area of a target building. The target building is a building that is the subject of this search. The target image is a candidate image including the target area identified based on the keywords after determining that the search requirement description information is in text form and performing keyword extraction on the search requirement description. For any candidate image, the keywords are used to construct prompt words for a large model. The candidate image and the prompt words are used as inputs of the large model to obtain a first output result or a second output result of the large model. The first output result indicates that the candidate image includes the target area, and the second output result indicates that the candidate image does not include the target area.
16. The device according to claim 15, wherein The first image set also includes: camera identifiers of cameras corresponding to the candidate images, wherein the cameras corresponding to any candidate image are cameras used to capture the candidate images, and each floor of the target building is provided with at least one camera; The retrieval results include: the cloud service determines the location information corresponding to the target image by querying the correspondence between each pre-saved camera identification and the corresponding location information, wherein any location information includes: the floor where the corresponding camera is located and the coordinate information of the corresponding camera on the floor.
17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 8.
19. A computer program product comprising a computer program / instructions, wherein when the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Image recognition combined with personal assistants for item recovery
US20200175302A1
Systems and methods for identifying a design template matching a search query
US20240311422A1