WiFi positioning method based on multi-modal knowledge base and related device
By employing a WiFi positioning method that integrates multimodal data fusion and dynamic knowledge updates, the problems of positioning accuracy and efficiency in dynamic 3D scenes and zero-sample scenarios are solved, achieving efficient and accurate object positioning.
Patent Information
- Application Number
- CN202510562576.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2026-01-09
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Traditional technologies struggle to efficiently integrate point cloud, image, text, and signal data in dynamic 3D scenes, and cannot accurately generate relevant descriptions in zero-sample scenarios, resulting in low positioning accuracy and efficiency.
By fusing multimodal data from point clouds, images, text, and WiFi signals, and combining dynamic knowledge updates and efficient retrieval and generation strategies, the CLIP Transformer model is used for multimodal feature alignment, and deep learning algorithms and image recognition technology are combined for object localization.
It significantly improves the accuracy and efficiency of positioning in complex environments, and is especially suitable for object positioning needs in dynamic 3D scenes.
Smart Images

Figure CN120086416B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer vision and multi-modal data processing, and particularly relates to a WiFi positioning method based on a multi-modal knowledge base and a related device. BACKGROUND
[0002] In recent years, multi-modal retrieval and generation (RAG) technology has made remarkable achievements in integrating text, vision and other data types.
[0003] However, traditional technology often only focuses on a single modality or a preset retrieval environment, and has difficulties in dealing with the following complex problems: 1) multi-modal fusion problem: in a dynamic 3D scene, it is extremely challenging to efficiently integrate point cloud, image, text and signal data. 2) Zero-shot scene generalization problem: when facing unseen data, accurate generation of related descriptions requires large-scale cross-modal alignment and context reasoning capabilities. SUMMARY
[0004] The purpose of the present application is to provide a WiFi positioning method based on a multi-modal knowledge base and a related device, which significantly improves the accuracy and efficiency of positioning in complex environments by fusing point cloud, image, text and WiFi signal multi-modal data, combining dynamic knowledge update and efficient retrieval generation strategy.
[0005] To achieve the above purpose, the present application provides the following solutions:
[0006] In a first aspect, the present application provides a WiFi positioning method based on a multi-modal knowledge base, comprising:
[0007] Obtaining item retrieval information submitted by a user; the item retrieval information includes a storage area of an item and an item feature description.
[0008] Based on a language model, performing text extraction on the item retrieval information to obtain a text feature of the item retrieval information.
[0009] According to the storage area of the item, determining a point cloud feature of the storage area; and performing multi-modal feature alignment on the point cloud feature and the text feature through a CLIP Transformer model to obtain an item feature description; the item feature description is stored in a multi-modal knowledge base.
[0010] Based on a WiFi signal strength indicator in the storage area of the item, measuring the strength of the WiFi signal in the storage area of the item, and determining the average signal strength and the fluctuation range of the signal strength of each WiFi access point in the storage area in the storage area.
[0011] obtain image information of the storage area based on a camera device of the storage area.
[0012] According to the item feature description, the average signal strength in the storage area, the fluctuation range of the signal strength, and the image information of the storage area, an item positioning analysis model is used to position the item in the item retrieval information submitted by the user; the item positioning analysis model is a prior probability analysis model introducing distance and area weight; the item positioning analysis model is used to perform semantic understanding on the item feature description by using a deep learning algorithm, determine the probability of the existence of the item in different areas in the storage area according to the average signal strength in the storage area and the fluctuation range of the signal strength, and identify the visual feature of the item by processing the image information of the storage area through an image recognition technology.
[0013] In a second aspect, the application provides a WiFi positioning device based on a multi-modal knowledge base, comprising:
[0014] The retrieval information acquisition module is configured to acquire item retrieval information submitted by a user; the item retrieval information includes a storage area of an item and an item feature description.
[0015] The text extraction module is configured to perform text extraction on the item retrieval information based on a language model to obtain text features of the item retrieval information.
[0016] The point cloud feature extraction module is configured to determine point cloud features of the storage area of the item according to the storage area of the item, and perform multi-modal feature alignment on the point cloud features and the text features through a CLIP Transformer model to obtain an item feature description; the item feature description is stored in a multi-modal knowledge base.
[0017] The WiFi strength measurement module is configured to measure the strength of a WiFi signal of the storage area of the item based on a WiFi signal strength indicator in the storage area of the item, and determine the average signal strength and the fluctuation range of the signal strength of each WiFi access point in the storage area in the storage area.
[0018] The image acquisition module is configured to obtain image information of the storage area based on a camera device of the storage area.
[0019] The positioning module is configured to position the article in the article search information submitted by the user based on an article positioning analysis model according to the article feature description, the average signal strength and the fluctuation range of the signal strength in the storage area, and the image information of the storage area; the article positioning analysis model is a prior probability analysis model introducing a distance and a region weight; the article positioning analysis model is configured to perform semantic understanding on the article feature description by using a deep learning algorithm, determine the probability of the existence of the article in different regions in the storage area according to the average signal strength and the fluctuation range of the signal strength in the storage area, and identify the visual feature of the article by processing the image information of the storage area by using an image recognition technology.
[0020] In a third aspect, the present application provides a computer device, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor executes the computer program to implement the WiFi positioning method based on the multi-modal knowledge base according to any one of the above.
[0021] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the WiFi positioning method based on the multi-modal knowledge base according to any one of the above.
[0022] In a fifth aspect, the present application provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the WiFi positioning method based on the multi-modal knowledge base according to any one of the above.
[0023] According to the embodiments provided in the present application, the following technical effects are disclosed:
[0024] The application provides a WiFi positioning method based on a multi-modal knowledge base and related devices. First, the user-submitted item retrieval information is obtained: the user submits retrieval information containing the item storage area and the item feature description, i.e., the item the user wants to find and the approximate location. Then, a language model is used to extract text from the item retrieval information to obtain text features. According to the storage area of the item, the point cloud features of the area are determined, i.e., the accurate three-dimensional spatial structure information of the storage area is obtained. A CLIP Transformer model is used to integrate visual and text information and enhance the understanding of item features. The WiFi signal strength indicator in the item storage area is used to measure the WiFi signal strength of the area, providing a wireless signal level reference for positioning. The average signal strength and signal strength fluctuation range of each WiFi access point in the storage area are determined. Image information is obtained based on the camera equipment in the storage area, and visual data is further used to assist the positioning process. Finally, the item positioning analysis model is used for positioning by combining the item feature description, the average signal strength, the signal strength fluctuation range, and the image information. The application ensures the efficiency and accuracy of item positioning, and is especially suitable for positioning needs in complex environments. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0026] Figure 1 An application environment diagram of a WiFi positioning method based on a multi-modal knowledge base in an embodiment of the present application.
[0027] Figure 2 A flowchart of a WiFi positioning method based on a multi-modal knowledge base provided by an embodiment of the present application.
[0028] Figure 3 An information retrieval diagram provided by an embodiment of the present application.
[0029] Figure 4 A storage area diagram provided by an embodiment of the present application.
[0030] Figure 5 A multi-modal knowledge picture dependency diagram provided by an embodiment of the present application.
[0031] Figure 6 A functional module diagram of a WiFi positioning device based on a multi-modal knowledge base provided by an embodiment of the present application.
[0032] Figure 7 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0033] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0034] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0035] The WiFi positioning method provided in this application embodiment can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on another server. Terminal 102 can send the item retrieval information to be processed to server 104. After receiving the item retrieval information, server 104 performs text extraction on the item retrieval information based on a language model to obtain the text features of the item retrieval information; determines the point cloud features of the storage area based on the storage area of the item; and aligns the point cloud features with the text features using a CLIP Transformer model to obtain an item feature description; the item feature description is stored in a multimodal knowledge base; measures the WiFi signal strength of the storage area based on a WiFi signal strength indicator, and determines the average signal strength and signal strength fluctuation range of each WiFi access point in the storage area; obtains image information of the storage area based on a camera device in the storage area; and locates the item in the item retrieval information submitted by the user based on the item feature description, the average signal strength and signal strength fluctuation range in the storage area, and the image information of the storage area, using an item location analysis model. Server 104 can then feed back the obtained item location to terminal 102. In addition, in some embodiments, the WiFi positioning method can also be implemented by the server 104 or the terminal 102 alone. For example, the terminal 102 can directly perform WiFi positioning on the item retrieval information to be processed, or the server 104 can obtain the item retrieval information to be processed from the data storage system and perform WiFi positioning on the item retrieval information to be processed.
[0036] The terminal 102 can be, but is not limited to, various desktop computers, notebook computers, smart phones, tablet computers, Internet of Things devices, and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle-mounted device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented by a standalone server or a server cluster composed of multiple servers, and can also be a cloud server.
[0037] In an exemplary embodiment, as shown in Figure 2 , a WiFi positioning method is provided, which is executed by a computer device, specifically by a terminal or a server, or by both a terminal and a server. In the embodiments of the present application, the method is applied to the server 104 in Figure 1 , and includes the following steps 201 to 206. Wherein:
[0038] Step 201: obtaining the item retrieval information submitted by the user; the item retrieval information includes the storage area of the item and the item feature description.
[0039] Step 202: based on the language model, text extraction is performed on the item retrieval information to obtain the text features of the item retrieval information.
[0040] Step 203: determining the point cloud features of the storage area according to the storage area of the item; and performing multi-modal feature alignment on the point cloud features and the text features through a CLIP Transformer model to obtain the item feature description; the item feature description is stored in a multi-modal knowledge base.
[0041] Step 204: based on the WiFi signal strength indicator in the item storage area, measuring the strength of the WiFi signal in the item storage area, and determining the average signal strength and the fluctuation range of the signal strength of each WiFi access point in the storage area.
[0042] Step 205: based on the camera device of the storage area, obtaining the image information of the storage area.
[0043] Step 206: Based on the item feature description, the average signal strength and the fluctuation range of the signal strength in the storage area, and the image information of the storage area, the item in the item search information submitted by the user is positioned based on an item positioning analysis model; the item positioning analysis model is an analysis model of prior probability with distance and area weight introduced; the item positioning analysis model is used for semantic understanding of the item feature description by using a deep learning algorithm, determination of the probability of existence of the item in different areas in the storage area according to the average signal strength and the fluctuation range of the signal strength in the storage area, and identification of the visual feature of the item by image recognition technology processing the image information of the storage area.
[0044] In some embodiments, when step 201 is performed, the following can be specifically implemented:
[0045] Obtaining item search information submitted by a user; the item search information includes a storage area of an item and an item feature description. For example, user A wants to find a water cup on the stairs. The item search information submitted by the user can be "What are the coordinates of the water cup on the stairs?" as shown in the following table. Figure 3
[0046] In some embodiments, when step 202 is performed, the following can be specifically implemented:
[0047] Using a trained BERT language model or a T5 language model to perform text extraction on the item search information to obtain a text feature of the item search information; the text feature is represented in a vector form.
[0048] Specifically, F t = LanguageEncoder .
[0049] Wherein, F t ∈Rdt represents a text feature vector, and Q is a text input.
[0050] In some embodiments, when step 203 is performed, the following can be specifically implemented:
[0051] Obtaining a three-dimensional model of a storage area; the three-dimensional model includes a plurality of point cloud data.
[0052] Using an efficient point cloud learning strategy, combining farthest point sampling and K- nearest neighbor technology, constructing a point cloud data block, and generating a point cloud feature based on the point cloud data block by using an MLP neural network.
[0053] Based on the trained CLIP Transformer model, the point cloud features and text features are fused to obtain the item feature description; the item feature description is represented in a multi-modal representation, , T task task embeddings, E pos position encodings, T p point cloud features.
[0054] Specifically, point cloud data is usually obtained from 3D scanning. In this embodiment, an Efficient Point Cloud Learning (EPCL) method is used to generate point cloud data blocks P through farthest point sampling (FPS) and K-Nearest Neighbors (KNN), and then process them through a multi-layer perceptron (MLP) to generate embeddings:
[0055] .
[0056] The point cloud features and text features are aligned through a CLIP Transformer model to form a multi-modal representation: Then, the text features and point cloud features are combined to align different types of features through a trained CLIP Transformer model (this model is frozen, meaning its parameters will not change during training), forming a unified multi-modal representation to ensure that the retrieval task can use the latest multi-modal information in real time. Task-specific embeddings (T task ) and position encodings (E pos ) are also added to help the system better understand the meaning of these features in a specific task.
[0057] Specifically, the item feature description can be in the following form:
[0058] "bssid" : "a", "x" : 9, "y" : 48, "channel" : 149; "name" : "r6-8", "top_left" : [1, 45], "bottom_right" : [17, 50], "aps" : ["a"]; "name" : "r6-6", "top_left" : [1, 91], "bottom_right" : [17, 98], "aps" : ["e"].
[0059] These are JSON-formatted data, commonly used in configuration files or data exchange. The specific field meanings are as follows:
[0060] 1) "bssid": "a" - represents the Basic Service Set Identifier (BSSID), which is a unique identifier for an access point in a wireless network; "x": represents some coordinate or numerical value; "y": represents a coordinate or numerical value; "channel": represents the channel number used by the wireless network.
[0061] 2) "name": "r6-6": represents a key-value pair indicating an attribute or object named "r6-6"; "top_left": [1, 91]: this key-value pair represents a coordinate point, where "top_left" may represent the top-left corner coordinate of a rectangle or region. The coordinate value is [1, 91], meaning that in a certain coordinate system, the x-coordinate is 1 and the y-coordinate is 91.
[0062] "bottom_right": [17, 98]: this key-value pair represents a coordinate point, where "bottom_right" may represent the bottom-right corner coordinate of a rectangle or region. The coordinate value is [17, 98], meaning that in the same coordinate system, the x-coordinate is 17 and the y-coordinate is 98.
[0063] "aps": ["e"]: this key-value pair represents an array, with the key "aps" and the value being an array containing a single element "e". This represents a certain attribute or identifier, such as a code that may represent a certain device, function, or state.
[0064] In some embodiments, before performing step 204, the following steps are also included:
[0065] Based on the obtained item feature description, a multi-modal query (such as text, image, point cloud, etc.) is performed, and an embedding vector v is generated using a feature extractor. The embedding vector v is stored in a vector database. During retrieval, the similarity between the query vector v and the stored vectors h in the database is calculated using weighted pooling.
[0066] The formula is as follows:
[0067] .
[0068] where v represents the embedding vector of the query, representing the multi-modal features of the query content; h represents the embedding vector stored in the database; w represents the position weight, used to emphasize the importance of matching at a certain position. Finally, the most matching result is returned according to the similarity score.
[0069] Before matching the optimal result, for each matching, the InfoNCE-based contrastive loss is used to optimize the retrieval performance, with the specific formula as follows:
[0070] .
[0071] where d + is the positive sample, D − is the negative sample set, and τ is the temperature parameter of the embedding space:
[0072] The retrieval strategy relies on the expressiveness of the embedding vectors v and h, while the loss optimization adjusts the embedding space through contrastive learning, so that: similar samples (e.g., v and positive sample h + ) are closer in the embedding space; dissimilar samples (e.g., v and negative sample h − ) are further apart in the embedding space. The optimized embedding space improves the effectiveness of the retrieval strategy, making the similarity calculation more accurate.
[0073] For the reliability of similarity measurement:
[0074] The retrieval strategy calculates similarity based on weighted pooling, and the loss optimization is based on this similarity to strengthen the discriminability of this measurement method through positive and negative sample comparison. Loss optimization helps the retrieval strategy to more reliably distinguish between high and low similarity samples.
[0075] For target consistency:
[0076] The goal of the retrieval strategy is to improve the accuracy of the retrieval results, while the goal of the loss optimization is to optimize the embedding vectors and similarity calculation methods used in the retrieval process. Both ultimately point to the improvement of retrieval performance.
[0077] Specifically, in the multi-modal knowledge retrieval enhancement. First, for a query containing multiple types of information (such as text, pictures, etc.), an embedding vector v is generated and stored in the vector database. When retrieving, the similarity of the retrieval results is calculated through weighted pooling, where w i is the position weight, h i is the embedding vector.
[0078] Next, the InfoNCE-based contrastive loss is used to optimize retrieval performance. Here, d + represents the positive sample, D - represents the negative sample set, and τ is the temperature parameter of the embedding space. The goal of optimization is to adjust the embedding space so that similar samples (e.g., v and positive sample h + ) are closer in the embedding space, while dissimilar samples (e.g., v and negative sample h - ) are further apart in the embedding space. This can improve the effectiveness of the retrieval strategy, making the similarity calculation more accurate.
[0079] In addition, the loss optimization also enhances the reliability of the similarity measure, strengthens the discriminability of the similarity measure method through the comparison of positive and negative samples, and helps the retrieval strategy to more reliably distinguish between high-similarity and low-similarity samples. Finally, the retrieval strategy and the loss optimization are both aimed at improving retrieval performance, and the goals are consistent.
[0080] It can be seen that the retrieval strategy performs similarity calculation based on the node embedding of the graph, and the dynamic updating mechanism ensures that the node embedding can be updated in real time to reflect the latest semantic information. The problem of outdated graph node information in traditional retrieval methods that cannot match the latest query is solved, and the real-time performance and accuracy of multi-modal retrieval are enhanced.
[0081] In some embodiments, when step 204 is performed, the following can be specifically implemented:
[0082] Assuming in an indoor environment, WiFi signal strength (RSSI value) can be measured to help positioning. These signal strength values are assumed to follow a mathematical rule called Gaussian distribution. To describe this distribution, the mean and variance of each WiFi communication positioning source within a specific area need to be known. Among them, such as Figure 4 The deployment of each WiFi signal device in different storage areas is shown. Among them, WiFi communication positioning sources 1, 2, and 3 are located on the first floor of the indoor area, and WiFi communication positioning sources 4 and 5 are located on the second floor of the indoor area.
[0083] In some embodiments, when step 205 is performed, the following can be specifically implemented:
[0084] By deploying a camera device in the specified storage area, image information in the storage area can be effectively obtained, thereby realizing real-time monitoring and management of the stored items; the camera device can be a camera carried on the robot.
[0085] In some embodiments, when step 206 is performed, the following can be specifically implemented:
[0086] Through a deep learning algorithm, the relevance of the item feature description and the image information of the storage area is determined.
[0087] Using the average signal strength and the fluctuation range of the signal strength, a machine learning technique is used to determine the signal strength analysis result; the signal strength analysis result is a mapping relationship between the signal strength and the item position.
[0088] According to the item feature description, the image information, and the signal strength analysis result, based on the prior probability and the posterior probability, the probability of the existence of items in different areas of the storage area is determined, and the Bayesian updating method is used to update the item position.
[0089] Wherein, in order to make the positioning more stable and reliable, the concept of prior probability is introduced. Prior probability considers the distance of the device from the last positioning position ( d i,last ) and the importance weight of the region ( w i ), The specific formula is as follows:
[0090] .
[0091] The importance weight can be understood as some regions are more important or more reliable than other regions. The smoothing factor ( ) is used to adjust the calculation method of prior probability to avoid the influence of extreme value.
[0092] Wherein, the calculation formula of posterior probability is:
[0093] .
[0094] In the formula, is the importance weight of the region, is the distance from the last positioning position, is the smoothing factor, M is the number of access points, and are the mean and variance of the access points in the region Ri; is the current RSSI value.
[0095] Specifically, according to the item feature description, image information and signal strength analysis result, based on the prior probability and the posterior probability, the probability of the existence of items in different regions in the storage area is determined, which specifically includes:
[0096] According to the formula , the probability of the existence of items in the region R i in the storage area is determined.
[0097] In some embodiments, after performing the item retrieval of the current user, the multi-modal knowledge base also needs to be updated, which is as follows:
[0098] Figure 5 The multi-modal knowledge picture represented in the dependency graph is updated by the knowledge graph (DMKG) through the following formula:
[0099] .
[0100] Wherein, controls the balance of new and old information, represents the node feature of the current time step. U t-1 is the graph embedding of the last time.
[0101] A relational graph attention network (RGAT) is used to generate context-dependent node embeddings:
[0102] .
[0103] where G" is a current context-dependent subgraph.
[0104] The optimization objective is to ensure that positive samples are close and negative samples are far away through a triple loss, and the specific formula is as follows:
[0105] .
[0106] wherein: p j , p k , p m is an anchor point, a positive sample, and a negative sample; delta is a boundary parameter.
[0107] Based on the same inventive concept, the embodiments of the present application also provide a WiFi positioning device for implementing the WiFi positioning method based on a multi-modal knowledge base. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more WiFi positioning device embodiments provided below can refer to the limitations of the WiFi positioning method in the foregoing, which will not be repeated here.
[0108] In one exemplary embodiment, as shown in Figure 6 a WiFi positioning method and device based on a multi-modal knowledge base are provided, comprising:
[0109] The retrieval information acquisition module 601 is configured to acquire item retrieval information submitted by a user; the item retrieval information includes a storage area of an item and an item feature description.
[0110] The text extraction module 602 is configured to perform text extraction on the item retrieval information based on a language model to obtain a text feature of the item retrieval information.
[0111] The point cloud feature extraction module 603 is configured to determine a point cloud feature of the storage area according to the storage area of the item, and perform multi-modal feature alignment on the point cloud feature and the text feature through a CLIP Transformer model to obtain an item feature description; the item feature description is stored in a multi-modal knowledge base.
[0112] The WiFi strength measurement module 604 is configured to measure the strength of the WiFi signal of the storage area of the article based on the WiFi signal strength indicator in the storage area of the article, and determine the average signal strength and the fluctuation range of the signal strength of each WiFi access point in the storage area in the storage area.
[0113] The image acquisition module 605 is configured to obtain image information of the storage area based on the camera device of the storage area.
[0114] The positioning module 606 is configured to position the article in the article retrieval information submitted by the user based on an article positioning analysis model according to the article feature description, the average signal strength and the fluctuation range of the signal strength in the storage area, and the image information of the storage area; the article positioning analysis model is a prior probability analysis model introducing a distance and a region weight; the article positioning analysis model is configured to perform semantic understanding on the article feature description by using a deep learning algorithm, determine the probability of existence of the article in different regions in the storage area according to the average signal strength and the fluctuation range of the signal strength in the storage area, and identify the visual feature of the article by processing the image information of the storage area through an image recognition technology.
[0115] In an exemplary embodiment, a computer device can be provided, which can be a server or a terminal, and an internal structure diagram thereof can be as shown in Figure 7 The computer device includes a processor, a memory, an input / output interface (I / O) and a communication interface. The processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The database of the computer device is configured to store WiFi positioning data. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement a WiFi positioning method based on a multi-modal knowledge base.
[0116] Those skilled in the art can understand that Figure 7 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. The specific computer device can include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0117] In an example embodiment, a computer device is also provided, including a memory and a processor, the memory storing a computer program, and the processor implementing the steps in the above method embodiments when executing the computer program.
[0118] In an example embodiment, a computer readable storage medium is provided, storing a computer program, and the computer program implementing the steps in the above method embodiments when executed by a processor.
[0119] In an example embodiment, a computer program product is provided, including a computer program, and the computer program implementing the steps in the above method embodiments when executed by a processor.
[0120] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0121] It can be understood by those skilled in the art that all or part of the processes in the above method embodiments can be completed by a computer program instructing related hardware, and the computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Any reference to memory, database or other medium used in the embodiments provided by the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0122] The database involved in each embodiment provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a blockchain, and the like, without being limited thereto. The processor involved in each embodiment provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, and the like, without being limited thereto.
[0123] In summary, the present application has the following technical effects:
[0124] 1) The present application integrates visual, language and WiFi data through real-time updating of a relational graph attention network (RGAT), and supports accurate retrieval and scene understanding.
[0125] 2) The present application supports real-time multi-modal processing and positioning: a lightweight visual-language model (MobileVLM) and a WiFi-based Bayesian positioning system are introduced to support real-time, multi-floor high-precision positioning.
[0126] 3) Update library of autonomous ability for complex tasks: can simulate to real verification (Sim-to-Real) for robots, and provide an effective semantic knowledge base for autonomous navigation and control tasks in dynamic environments.
[0127] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, but as long as the combinations of the technical features do not contradict, they should be considered as the scope of the present application.
[0128] The principles and implementation modes of the present application are described by applying specific examples in this paper, and the above embodiment descriptions are only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In summary, the content of the present application should not be understood as a limitation.
Claims
1. A WiFi positioning method based on a multi-modal knowledge base, characterized in that, The WiFi positioning method based on the multi-modal knowledge base comprises: Obtaining the item retrieval information submitted by the user; the item retrieval information includes the storage area of the item and the item feature description; Based on the language model, the text feature of the item retrieval information is obtained by text extraction; According to the storage area of the item, the point cloud feature of the storage area is determined, and the point cloud feature and the text feature are aligned through the CLIP Transformer model to obtain the item feature description; the item feature description is stored in the multi-modal knowledge base; According to the storage area of the item, the point cloud feature of the storage area is determined, and the point cloud feature and the text feature are aligned through the CLIP Transformer model to obtain the item feature description, specifically including: Obtaining the three-dimensional model of the storage area; the three-dimensional model includes a plurality of point cloud data; Using an efficient point cloud learning strategy, combining the farthest point sampling and K- nearest neighbor technology, constructing a point cloud data block, and generating a point cloud feature based on the point cloud data block using an MLP neural network; Based on the trained CLIP Transformer model, the point cloud features and the text features are fused to obtain an item feature description; the item feature description is represented in a multi-modal representation manner, , T task for task embedding, E pos for position encoding, T p for point cloud features; The specific form of the item feature description is as follows: bssid, x, y; bssid is the unique identifier of the access point in the wireless network; x represents a certain coordinate or value; y represents the coordinate or value; Based on the WiFi signal strength indicator in the storage area of the item, the strength of the WiFi signal in the storage area of the item is measured, and the average signal strength and the fluctuation range of the signal strength of each WiFi access point in the storage area are determined; Based on the camera equipment in the storage area, the image information of the storage area is obtained; According to the item feature description, the average signal strength and the fluctuation range of the signal strength in the storage area, and the image information of the storage area, the item in the item retrieval information submitted by the user is positioned based on the item positioning analysis model; the item positioning analysis model is an analysis model introducing distance and regional weight prior probability; the item positioning analysis model is used to perform semantic understanding on the item feature description by using a deep learning algorithm, determine the probability of the existence of the item in different regions in the storage area according to the average signal strength and the fluctuation range of the signal strength in the storage area, and identify the visual features of the item by processing the image information of the storage area through image recognition technology; According to the item feature description, the average signal strength and the fluctuation range of the signal strength in the storage area, and the image information of the storage area, the item in the item retrieval information submitted by the user is positioned based on the item positioning analysis model, specifically including: Determine the correlation between the item feature description and the storage area image information by using a deep learning algorithm; Using the average signal strength and the fluctuation range of the signal strength, a signal strength analysis result is determined by using a machine learning technology; the signal strength analysis result is the mapping relationship between the signal strength and the item position; According to the item feature description, image information and signal strength analysis result, the probability of existence of the item in different regions in the storage region is determined based on the prior probability and the posterior probability, and the item position is updated by using the Bayesian updating method; The calculation formula of the prior probability is: ; The calculation formula of the posterior probability is: ; wherein, is the importance weight of the zone, is the distance from the last position fix, is the smoothing factor, M is the number of access points, and is the mean and variance of the access points in the zone R i is the current RSSI value; According to the item feature description, image information and signal strength analysis result, the probability of existence of the item in different regions in the storage region is determined based on the prior probability and the posterior probability, and the item position is updated by using the Bayesian updating method; According to the formula determining the probability of the presence of an item in different areas within the storage area; wherein R i and R k are regions i and regions k . 2.The WiFi positioning method based on multi-modal knowledge base according to claim 1, characterized in that, The text feature of the item retrieval information is obtained by performing text extraction on the item retrieval information based on a language model, and the text feature of the item retrieval information is obtained by performing text extraction on the item retrieval information based on a language model, and the text feature of the item retrieval information is obtained by performing text extraction on the item retrieval information based on a language model. The WiFi positioning device comprises: 3.A multi-modal knowledge base based WiFi positioning device for implementing the multi-modal knowledge base based WiFi positioning method of claim 1, characterized in that, The retrieval information acquisition module is configured to acquire the item retrieval information submitted by the user, and the item retrieval information comprises the storage region of the item and the item feature description. The text extraction module is configured to perform text extraction on the item retrieval information based on a language model to obtain the text feature of the item retrieval information. The point cloud feature extraction module is configured to determine the point cloud feature of the storage region according to the storage region of the item, and perform multi-modal feature alignment on the point cloud feature and the text feature through a CLIP Transformer model to obtain the item feature description. The WiFi strength measurement module is configured to measure the strength of the WiFi signal in the storage region of the item based on the WiFi signal strength indicator in the storage region of the item, and determine the average signal strength and the fluctuation range of the signal strength of each WiFi access point in the storage region. The image acquisition module is configured to obtain the image information of the storage region based on the camera device of the storage region. The positioning module is configured to position the item in the item retrieval information submitted by the user based on an item positioning analysis model according to the item feature description, the average signal strength and the fluctuation range of the signal strength in the storage region, and the image information of the storage region. The memory, the processor and the computer program stored in the memory and executable on the processor are characterized in that the processor executes the computer program to implement the WiFi positioning method based on the multi-modal knowledge base according to any one of claims 1-2.
4. A computer device comprising: The computer program is executed by the processor to implement the WiFi positioning method based on the multi-modal knowledge base according to any one of claims 1-2.
5. A computer-readable storage medium having stored thereon a computer program, characterized in that, 6. A computer program product comprising a computer program, characterized in that, The computer program, when executed by a processor, implements the method of claim 1-2 for WiFi positioning based on a multi-modal knowledge base.
Citation Information
Patent Citations
Positioning system for unmanned supermarket based on wifi technology
CN115426619A
Position determination method and device and storage medium
CN117460044A
Embedded intelligent visual language large model knowledge base construction and application method, equipment, medium and product
CN119476463A