Cross-view-angle positioning prompt accurate coding method based on perception large model knowledge distillation
By introducing knowledge distillation technology from a large perceptual model into the cross-view localization model, the accuracy problem caused by semantic ambiguity in the localization model is solved, and a higher accuracy localization effect is achieved.
Patent Information
- Application Number
- CN202510886997.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-10-21
AI Technical Summary
Existing cross-view positioning models suffer from semantic ambiguity in point cue encoding methods, resulting in limited positioning accuracy and difficulty in accurately identifying users' location cue requests.
A knowledge distillation method based on the perceptual big model is adopted to distill the semantic understanding knowledge in the perceptual big model (SAM) into the localization model. Through mask-level and feature-level knowledge distillation, the localization model's ability to accurately understand user prompts is improved.
It improves the positioning accuracy and robustness of the interactive cross-view positioning model, enabling it to more accurately identify user location prompts and enhance the consistency of positioning results.
Smart Images

Figure CN120823264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a positioning hint encoding method, relates to the field of geographic positioning, and specifically to a cross-view positioning hint accurate encoding method based on knowledge distillation of a large perceptual model. Background Art
[0002] Cross-view positioning refers to identifying and locating objects in satellite images under cross-view conditions, such as drone-satellite and ground-satellite, by fusing visual information from different perspectives to obtain the geographic coordinates of the object to be located. Cross-view positioning has important application value in many fields such as navigation, monitoring, and search and rescue. Of course, the cross-view positioning model is generally an interactive model. The user needs to first point out the object to be located in the query image (such as a drone or ground image) by clicking. The model then returns the coordinates of the object the user wants to locate on the satellite image based on the user's point prompt, combined with the information of the query image and the reference image (such as a satellite image).
[0003] To help the model understand user click cues, existing point cue encoding methods use Euclidean distance to convert user clicks on an image into a mask representing the target object's location on the query image, allowing the model to understand which object it is targeting. Specifically, this Euclidean distance-based encoding method first locates the center point of the target object on the image. Then, by calculating the Euclidean distance, it generates a grayscale image with the center point in white and the outer circle diffused using a quadratic function, creating a gradient from white to black.
[0004] While point cue encoding methods based on Euclidean distance have achieved some success, their limitations are also significant. The primary issue is that when using Euclidean distance to represent an object, the resulting grayscale image cannot accurately depict the object's true outline, but only indicates its approximate range and location. This leads to semantic ambiguity in point cue encoding. Even if the points of an object are manually labeled, the object's spatial layout and shape cannot be accurately represented, which affects the model's understanding of user cues and reduces the model's positioning accuracy.
[0005] Currently, the positioning accuracy of the model is limited using the click prompt method. The semantic ambiguity of the current point prompt encoding method reduces the positioning accuracy of the positioning model. Specifically, the existing point prompt encoding method uses point prompt encoding based on Euclidean distance. This encoding method can only calibrate the approximate position and approximate range of the object to be located on the query image. This makes the prompt for the object to be located semantically ambiguous when the query image content is complex. When faced with ambiguous prompts, the positioning model finds it difficult to accurately locate the target that meets the user's prompt requirements, which in turn reduces the model's positioning accuracy. Summary of the Invention
[0006] In order to solve the problems existing in the background technology, the present invention provides a method for accurately encoding cross-view positioning prompts based on knowledge distillation of a perceptual large model. Taking into account the problem that the accuracy of the positioning model is limited due to the semantic ambiguity of the existing point prompt encoding method in the current interactive cross-view positioning model, this method utilizes the semantic understanding knowledge of the point prompt in the existing perceptual large model SAM (Segment Anything Model) to construct a method for accurately encoding cross-view positioning prompts based on knowledge distillation of a perceptual large model. This method replaces the prompt encoder in the existing interactive cross-view positioning model, so that the model obtains an accurate understanding of the user prompts, thereby introducing accurate user prompt information into the positioning model. Compared with the existing method, the present method can effectively improve the positioning accuracy of the positioning model for the positioning target desired by the user.
[0007] The technical solution adopted in the present invention is:
[0008] The cross-view positioning hint accurate encoding method based on perceptual large model knowledge distillation of the present invention includes:
[0009] Step 1) Set up a teacher model T and a student model S based on a perception large model, encoding operation and positioning aggregation module, train the teacher model T and the student model S through several geographic location images marked with location prompt points, and use the knowledge distillation method to perform knowledge distillation on the teacher model T and the student model S. At the same time, construct a total knowledge distillation loss function until the total knowledge distillation loss function converges, obtain the positioning aggregation module in the trained student model S, and replace the aggregation module of the original positioning model to obtain the trained positioning model.
[0010] Step 2) The query image marked with the location cue point to be queried and the cross-view reference image are input into the trained positioning model for processing, so as to obtain the coordinate positioning boundary box of the location cue point to be queried in the reference image and then display it on the display screen, realizing accurate encoding of cross-view positioning cues.
[0011] In the step 1), the teacher model T includes a perceptual large model SAM, a first convolutional coding and a first positioning aggregation module connected in sequence, and the student model S includes a Euclidean distance coding, a second convolutional coding and a second positioning aggregation module connected in sequence. The first and second positioning aggregation modules are both aggregation modules of the original positioning model; the positioning model adopts an interactive cross-perspective geographic positioning model.
[0012] In the step 1), the knowledge distillation method includes mask-level knowledge distillation and feature-level knowledge distillation; when the teacher model T is trained, the first query information aggregation feature F is obtained. T and the first query information aggregated word TT , obtain the second query information aggregation feature F when the student model S is trained S and the second query information aggregated word T S , based on the second query information aggregation feature F S Constructing mask distillation loss for mask-level knowledge distillation Aggregate features F based on the first query information T , the first query information aggregated word T T , the second query information aggregation feature F S and the second query information aggregated word T S Constructing feature distillation loss for feature-level knowledge distillation Mask-based distillation loss and feature distillation loss Constructing the total knowledge distillation loss function
[0013] In the step 1), for each image with a position prompt point P q The geographic location image Iq and the location prompt point P are used when training the teacher model T. q The information is input into the perception model SAM of the teacher model T for processing to obtain the first object-level mask hint m SAM , position prompt point P q The information includes the position hint point P q The pixel coordinates and position category labels at the location; then the first object level mask prompt m SAM The first hint code is obtained by encoding the learnable parameters in the convolutional layer of the first convolutional code, and then input into the first positioning aggregation module for processing to obtain the query information aggregation feature F containing the semantic understanding prior knowledge T and query information aggregation word T T .
[0014] In the step 1), for each location hint point P marked in the location image Iq, q , when the student model S is trained, the position prompt point P q The information of is input into the student model S and the second object level mask hint m is obtained by Euclidean distance encoding. o , Euclidean distance encoding is a point hint encoding method based on Euclidean distance; then the second object level mask hint m o The second hint code is obtained by encoding the learnable parameters in the convolutional layer of the second convolutional code, and then input into the second positioning aggregation module for processing to obtain the second query information aggregation feature F that does not contain semantic understanding prior knowledge S and the second query information aggregated word T S .
[0015] The mask distillation loss The details are as follows:
[0016]
[0017] Among them, σ() is the Sigmoid activation function and conv() is the convolutional layer.
[0018] The characteristic distillation loss The details are as follows:
[0019]
[0020] Among them, σ() is the Sigmoid activation function, and mean() is the channel-wise average.
[0021] The total knowledge distillation loss function The details are as follows:
[0022]
[0023] Wherein, α and β are the first and second dynamic weight parameters respectively.
[0024] The electronic device of the present invention comprises: a memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method described above.
[0025] The computer-readable storage medium of the present invention stores program data thereon, and when the program data is executed by a processor, the method described above is implemented.
[0026] This method leverages the interactive semantic segmentation prior knowledge contained in the existing large-scale perceptual model (SAM). Through knowledge distillation, it distills the SAM's ability to accurately understand semantic cues into the cue encoding portion of the original positioning model, guiding the positioning model to accurately understand semantic cues and thus achieve precise cross-viewpoint geolocation. By precisely encoding user cues, this method reduces the semantic ambiguity of the cue encoding. The large-scale perceptual model (SAM) itself possesses rich knowledge of point cue semantic understanding, thereby improving the positioning model's accuracy.
[0027] The beneficial effects of the present invention are:
[0028] This method takes into account the problem of low positioning accuracy of the current interactive cross-view positioning model due to the semantic ambiguity of user point prompts during interactive positioning, as well as the problem of limited positioning model accuracy due to the semantic ambiguity of the existing point prompt encoding method. From the perspective of improving the accuracy of encoded user prompts, the semantic understanding knowledge of the perceptual large model SAM is used to realize contour-level prompt encoding of the target to be located, and the semantic understanding knowledge of the perceptual large model SAM is transferred into the positioning model through knowledge distillation. This avoids the huge computational overhead brought by the use of the perceptual large model SAM while improving the accuracy and robustness of the interactive cross-view geolocation model. It can effectively improve the positioning accuracy of the existing interactive cross-view geolocation, and can improve the semantic understanding ability of user prompts, thereby providing positioning results that are more in line with user intentions. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 Schematic diagram of model knowledge distillation training of the method of the present invention;
[0030] Figure 2 This is a comparison diagram of the positioning effects of the original interactive positioning model in the embodiment of the present invention and the interactive positioning model using the present method for prompt coding. DETAILED DESCRIPTION
[0031] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0032] This paper provides a method for accurately encoding cross-viewpoint location cues based on knowledge distillation from a perceptual macromodel. This method embeds the semantic understanding prior knowledge contained in the existing perceptual macromodel (SAM) at the cue encoder stage of an existing interactive cross-viewpoint geolocation model. This prior knowledge from the macromodel improves the location model's ability to understand user point cues. The specific steps are as follows:
[0033] Step 1) Set up a teacher model T and a student model S based on a perception large model, encoding operation and positioning aggregation module, train the teacher model T and the student model S through several geographic location images marked with location prompt points, and use the knowledge distillation method to perform knowledge distillation on the teacher model T and the student model S. At the same time, construct a total knowledge distillation loss function until the total knowledge distillation loss function converges, obtain the positioning aggregation module in the trained student model S, and replace the aggregation module of the original positioning model to obtain the trained positioning model.
[0034] like Figure 1As shown, the teacher model T includes a perceptual large model SAM, a first convolutional coding, and a first positioning aggregation module connected in sequence, and the student model S includes a Euclidean distance coding, a second convolutional coding, and a second positioning aggregation module connected in sequence. The first and second positioning aggregation modules are both aggregation modules of the original positioning model; the positioning model adopts an interactive cross-view geolocation model, and the aggregation module is the prompt encoder part of the interactive cross-view geolocation model. The interactive cross-view geolocation model specifically adopts the DetGeo model. The knowledge distillation method includes mask-level knowledge distillation and feature-level knowledge distillation; when the teacher model T is trained, the first query information aggregation feature F is obtained. T and the first query information aggregated word T T , obtain the second query information aggregation feature F when the student model S is trained S and the second query information aggregated word T S , based on the second query information aggregation feature F S Constructing mask distillation loss for mask-level knowledge distillation Aggregate features F based on the first query information T , the first query information aggregated word T T , the second query information aggregation feature F S and the second query information aggregated word T S Constructing feature distillation loss for feature-level knowledge distillation Mask-based distillation loss and feature distillation loss Constructing the total knowledge distillation loss function
[0035] Mask-level knowledge distillation distills the semantic understanding prior knowledge in the perception model SAM from the teacher model T into the second positioning aggregation module of the student model S. Specifically, the perception model SAM is embedded in the second positioning aggregation module of the student model S by distilling the semantic understanding prior knowledge in the teacher model T into the second positioning aggregation module of the student model S. q The coordinates and location category labels (such as stores, schools, etc.) of the object are processed to generate the first object-level mask hint m containing the semantic category. SAM , the mask implicitly perceives the semantic segmentation ability of the object obtained by the pre-training of the large model SAM, that is, the semantic understanding prior knowledge, and through the mask distillation loss The student model S learns the mask generation logic; the feature-level knowledge distillation transfers the precise semantic feature encoding capability obtained by the perceptual large model SAM in the teacher model T to the second positioning aggregation module of the student model S, which is specifically manifested as follows: the teacher model T uses the first convolutional encoding with learnable parameters to encode the first object-level mask hint m SAM Processing is performed and combined with the positioning aggregation module to generate the first query information aggregation feature F containing semantic features such as object structure and category T and the first query information aggregated word T TThis process reflects the ability of the perceptual large model SAM to accurately encode semantic features, and the feature distillation loss Aggregate features F by constraining the second query information of the student model S S and the second query information aggregated word T S Aggregate feature F with the first query information T , the first query information aggregated word T T The difference between the two is used to realize the transmission of the coding capability.
[0036] In specific implementation, for each image with a position hint point P q The geographic location image Iq and the location prompt point P are used when training the teacher model T. q The information is input into the perception model SAM of the teacher model T for processing to obtain the first object-level mask hint m SAM , position prompt point P q The information includes the position hint point P q The pixel coordinates and location category labels at the current location are as follows: q The name of the building at the location, such as the name of the store, school, etc.; then the first object level mask prompt m SAM The first hint code is obtained by encoding the learnable parameters in the convolutional layer of the first convolutional code, and then input into the first positioning aggregation module for processing to obtain the query information aggregation feature F containing the semantic understanding prior knowledge T and query information aggregation word T T .
[0037] In specific implementation, for each location hint point P marked in the location image Iq, q , when the student model S is trained, the position prompt point P q The information of is input into the student model S and the second object level mask hint m is obtained by Euclidean distance encoding. o , Euclidean distance encoding is a point hint encoding method based on Euclidean distance; then the second object level mask hint m o The second hint code is obtained by encoding the learnable parameters in the convolutional layer of the second convolutional code, and then input into the second positioning aggregation module for processing to obtain the second query information aggregation feature F that does not contain semantic understanding prior knowledge S and the second query information aggregated word T S .
[0038] Mask Distillation Loss The details are as follows:
[0039]
[0040] Among them, σ() is the Sigmoid activation function and conv() is the convolutional layer.
[0041] Feature Distillation Loss The details are as follows:
[0042]
[0043] Among them, σ() is the Sigmoid activation function, and mean() is the channel-wise average.
[0044] Total knowledge distillation loss function The details are as follows:
[0045]
[0046] Wherein, α and β are the first and second dynamic weight parameters respectively.
[0047] Step 2) The query image marked with the location cue point to be queried and the cross-view reference image are input into the trained positioning model for processing, so as to obtain the coordinate positioning boundary box of the location cue point to be queried in the reference image and then display it on the display screen, realizing accurate encoding of cross-view positioning cues.
[0048] After the teacher model T and the student model S are trained using the method of the present invention, the student model S containing the knowledge of the perception large model SAM is used to obtain the interactive cross-view geolocation model that is ultimately used for actual reasoning.
[0049] like Figure 2 As shown, the interactive cross-perspective geographic positioning model DetGeo model is used as a display example, which shows the positioning effect of the basic positioning model on the object to be located, as well as the positioning effect of the interactive cross-perspective geographic positioning model DetGeo model enhanced after applying the method of the present invention on the object to be located. It can be seen that the positioning result after applying the method of the present invention is more in line with the user's prompts.
[0050] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems or computer program products. Therefore, the application can adopt the form of a complete hardware embodiment, a complete software embodiment or an embodiment in combination with software and hardware. Moreover, the application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) that contain computer-usable program code. The scheme in the embodiments of the present application can be implemented in various computer languages. The application is described according to the flow chart of the method, system and computer program product of the embodiments of the present application.
[0051] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concepts. Therefore, the present invention is intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
[0052] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the equivalent technology of the present invention, the present application is intended to include these modifications and variations.
Claims
1. A cross-view positioning cue accurate encoding method based on knowledge distillation of a perceptual large model, characterized by: include: Step 1) Setting up a teacher model T and a student model S based on a large perception model, encoding operation, and positioning aggregation module, training the teacher model T and the student model S through a number of geographic location images marked with location prompt points, and using the knowledge distillation method to perform knowledge distillation on the teacher model T and the student model S. At the same time, constructing a total knowledge distillation loss function until the total knowledge distillation loss function converges, obtaining the positioning aggregation module in the trained student model S, and replacing the aggregation module of the original positioning model to obtain the trained positioning model; Step 2) The query image marked with the location prompt point to be queried and the cross-view reference image are input into the trained positioning model for processing, so as to obtain the coordinate positioning of the location prompt point to be queried in the reference image and then display it on the display screen, realizing accurate encoding of cross-view positioning prompts.
2. The cross-view positioning cue precise encoding method based on perceptual large model knowledge distillation according to claim 1 is characterized by: In the step 1), the teacher model T includes a perceptual large model SAM, a first convolutional coding and a first positioning aggregation module connected in sequence, and the student model S includes a Euclidean distance coding, a second convolutional coding and a second positioning aggregation module connected in sequence. The first and second positioning aggregation modules are both aggregation modules of the original positioning model; the positioning model adopts an interactive cross-perspective geographic positioning model.
3. The cross-view positioning cue precise encoding method based on perceptual large model knowledge distillation according to claim 2 is characterized by: In the step 1), the knowledge distillation method includes mask-level knowledge distillation and feature-level knowledge distillation; when the teacher model T is trained, the first query information aggregation feature F is obtained. T and the first query information aggregated word T T , obtain the second query information aggregation feature F when the student model S is trained S and the second query information aggregated word T S , based on the second query information aggregation feature F S Constructing mask distillation loss for mask-level knowledge distillation Aggregate features F based on the first query information T , the first query information aggregated word T T , the second query information aggregation feature F S and the second query information aggregated word T S Constructing feature distillation loss for feature-level knowledge distillation Mask-based distillation loss and feature distillation loss Constructing the total knowledge distillation loss function 4. The cross-view positioning cue precise encoding method based on perceptual large model knowledge distillation according to claim 3 is characterized by: In the step 1), for each image with a position prompt point P q The geographic location image Iq and the location prompt point P are used when training the teacher model T. q The information is input into the perception model SAM of the teacher model T for processing to obtain the first object-level mask hint m SAM , position prompt point P q The information includes the position hint point P q The pixel coordinates and position category labels at the location; then the first object level mask prompt m SAM After encoding through the first convolutional coding, the first prompt code is obtained, which is then input into the first positioning aggregation module for processing to obtain the query information aggregation feature F containing the semantic understanding prior knowledge T and query information aggregation word T T .
5. The cross-view positioning cue precise encoding method based on perceptual large model knowledge distillation according to claim 3 is characterized by: In the step 1), for each location hint point P marked in the location image Iq, q , when the student model S is trained, the position prompt point P q The information of is input into the student model S and the second object level mask hint m is obtained by Euclidean distance encoding. o , and then the second object level mask hint m o After encoding through the second convolutional coding, the second prompt code is obtained, which is then input into the second positioning aggregation module for processing to obtain the second query information aggregation feature F that does not contain semantic understanding prior knowledge S and the second query information aggregated word T S .
6. The cross-view positioning cue precise encoding method based on perceptual large model knowledge distillation according to claim 3 is characterized by: The mask distillation loss The details are as follows: Among them, σ() is the Sigmoid activation function and conv() is the convolutional layer.
7. The cross-view positioning cue precise encoding method based on perceptual large model knowledge distillation according to claim 3 is characterized by: The characteristic distillation loss The details are as follows: Among them, σ() is the Sigmoid activation function, and mean() is the channel-wise average.
8. The cross-view positioning cue precise encoding method based on perceptual large model knowledge distillation according to claim 3 is characterized by: The total knowledge distillation loss function The details are as follows: Wherein, α and β are the first and second dynamic weight parameters respectively.
9. An electronic device, characterized in that: include: A memory and a processor coupled to each other, wherein the memory stores program data, and the processor calls the program data to execute the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having program data stored thereon, characterized in that: When the program data is executed by a processor, the method according to any one of claims 1 to 8 is implemented.