Relocation method and device

By generating pose semantic descriptors and filtering target keyframe images based on preset constraints, the problem of low relocation accuracy of monocular visual SLAM system is solved, and the accuracy of relocation is significantly improved.

CN119942053APending Publication Date: 2025-05-06NEUSOFT REACH AUTOMOBILE TECH (SHENYANG) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411803945.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-09
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

Existing monocular vision SLAM systems have low accuracy during repositioning.

Method used

By obtaining the semantic matrix and pose matrix of each frame image, a pose semantic descriptor is generated, and the target keyframe image is determined based on the descriptor and preset constraints, and finally repositioning is performed with the historical keyframe image.

Benefits of technology

The accuracy of relocation is improved, and the target keyframe images are accurately filtered by combining semantic information and pose information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942053A_ABST
    Figure CN119942053A_ABST
Patent Text Reader

Abstract

The invention discloses a repositioning method and device. The method comprises the following steps: acquiring a semantic matrix and a pose matrix corresponding to each frame of image; based on the semantic matrix and the pose matrix, generating a pose semantic descriptor; based on the pose semantic descriptor and preset constraint conditions, determining a target key frame image, the preset constraint conditions including a preset pose constraint condition, a preset category constraint condition and a preset frame constraint condition which are executed in sequence; and performing relocation based on the target key frame image and the historical key frame image. According to the method, the semantic matrix and the pose matrix are combined to generate the pose semantic descriptor, so that the target key frame image is screened through the pose semantic descriptor and triple preset constraint conditions, the accurate target key frame image is obtained, and the relocation accuracy is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of positioning and navigation technology, and in particular to a repositioning method and device. Background Art

[0002] In the field of robot positioning and navigation, simultaneous localization and mapping (SLAM) technology is an important method to achieve autonomous robot navigation and environmental understanding. Especially in the fields of mobile robots, drones, and autonomous driving, SLAM technology can help robots achieve positioning and path planning in unknown or dynamic environments.

[0003] Monocular vision SLAM is a common SLAM implementation method, which relies only on image data collected by a camera to calculate the robot's position information by incrementally building maps and estimating poses frame by frame. Although existing monocular SLAM systems have adopted a variety of improvement methods, such as introducing deep learning for target detection or semantic segmentation.

[0004] However, the above method has the problem of low repositioning accuracy. Summary of the invention

[0005] The main purpose of the present application is to provide a repositioning method and device to solve the problem of low repositioning accuracy in the prior art.

[0006] In order to achieve the above objectives, in a first aspect, the present application provides a relocation method, comprising:

[0007] Get the semantic matrix and pose matrix corresponding to each frame image;

[0008] Generate a pose semantic descriptor based on the semantic matrix and the pose matrix;

[0009] Determine a target key frame image based on the pose semantic descriptor and preset constraints, wherein the preset constraints include preset pose constraints, preset category constraints, and preset border constraints that are executed in sequence;

[0010] Relocalization is performed based on the target keyframe image and the historical keyframe images.

[0011] In a possible implementation, obtaining a semantic matrix corresponding to each frame image includes:

[0012] Get the number of objects, object category labels, and two-dimensional bounding box locations in each frame image;

[0013] The semantic matrix of each frame image is generated based on the number of objects in each frame image, the object category label and the two-dimensional bounding box position.

[0014] In a possible implementation, obtaining the pose matrix corresponding to each frame image includes:

[0015] Get the rotation matrix, scaling factor, and displacement of the camera in the world coordinate system;

[0016] The pose matrix of each frame image is generated by the rotation matrix, scaling factor and the displacement of the camera in the world coordinate system.

[0017] In a possible implementation, a pose semantic descriptor is generated based on a semantic matrix and a pose matrix, including:

[0018] Get the timestamp of each frame image;

[0019] Based on the semantic matrix and pose matrix corresponding to each frame image, as well as the timestamp of each frame image, a pose semantic descriptor is generated.

[0020] In a possible implementation, a pose semantic descriptor is generated based on the semantic matrix and pose matrix corresponding to each frame image and the timestamp of each frame image, including:

[0021] Extract the first column data of the semantic matrix corresponding to each frame image, and form a target category set based on the first column data of the semantic matrix;

[0022] Extract the last four columns of data of the semantic matrix corresponding to each frame image, and form a two-dimensional bounding box set from the last four columns of data of the semantic matrix;

[0023] The pose semantic descriptor is generated by the timestamp of each frame image, the set of target categories, the set of two-dimensional bounding boxes and the pose matrix.

[0024] In a possible implementation, determining a target key frame image based on a pose semantic descriptor and preset constraints includes:

[0025] Determine a first key frame image based on the pose semantic descriptor and preset pose constraints;

[0026] Determine a second key frame image based on the pose semantic descriptor, the preset category constraint and the first key frame image;

[0027] Based on the pose semantic descriptor, the preset bounding box constraint and the second key frame image, a target key frame image is determined.

[0028] In a possible implementation, determining a first key frame image based on a pose semantic descriptor and a preset pose constraint condition includes:

[0029] Extract the pose matrix in the pose semantic descriptor, and calculate the geometric similarity between the historical key frame image and the current key frame image based on the pose matrix;

[0030] A key frame image whose geometric similarity satisfies a preset position constraint condition is selected from the current key frame image as the first key frame image.

[0031] In a possible implementation, determining the second key frame image based on the pose semantic descriptor, the preset category constraint condition and the first key frame image includes:

[0032] Extract the target category set in the pose semantic descriptor, and calculate the category label similarity between the historical key frame image and the current key frame image based on the target category set;

[0033] A key frame image whose category label similarity satisfies a preset category constraint condition is selected from the first key frame image as the second key frame image.

[0034] In a possible implementation, determining a target key frame image based on a pose semantic descriptor, a preset bounding box constraint, and a second key frame image includes:

[0035] Extracting a two-dimensional bounding box set from the pose semantic descriptor, and calculating the overlap of the bounding boxes between the historical key frame image and the current key frame image based on the two-dimensional bounding box set;

[0036] A key frame image whose overlap degree satisfies a preset border constraint condition is selected from the second key frame image as a target key frame image.

[0037] In a second aspect, an embodiment of the present invention provides a relocation device, comprising:

[0038] An acquisition module is used to obtain the semantic matrix and posture matrix corresponding to each frame image;

[0039] A descriptor generation module, used for generating a pose semantic descriptor based on a semantic matrix and a pose matrix;

[0040] A target key frame determination module is used to determine a target key frame image based on a pose semantic descriptor and preset constraints, wherein the preset constraints include a preset pose constraint, a preset category constraint, and a preset border constraint that are executed in sequence;

[0041] The relocalization module is used to perform relocalization based on the target key frame image and the historical key frame images.

[0042] In a third aspect, an embodiment of the present invention provides a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above relocation methods when executing the computer program.

[0043] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above relocation methods are implemented.

[0044] The embodiment of the present invention provides a repositioning method and device, including: first obtaining the semantic matrix and pose matrix corresponding to each frame image, then generating a pose semantic descriptor based on the semantic matrix and pose matrix, and then determining the target key frame image based on the pose semantic descriptor and preset constraints, wherein the preset constraints include preset pose constraints, preset category constraints, and preset border constraints that are executed in sequence, and finally performing repositioning based on the target key frame image and the historical key frame image. The present application combines the semantic matrix and the pose matrix to generate a pose semantic descriptor, thereby screening the target key frame image through the pose semantic descriptor and the triple preset constraints, obtaining an accurate target key frame image, and thus improving the accuracy of repositioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings constituting a part of this application are used to provide a further understanding of this application, so that other features, purposes and advantages of this application become more obvious. The schematic embodiment drawings and their descriptions of this application are used to explain this application and do not constitute an improper limitation on this application. In the drawings:

[0046] Figure 1 is a flow chart of an implementation of a relocation method provided by an embodiment of the present invention;

[0047] Figure 2 is a structural schematic diagram of a relocation device provided by an embodiment of the present invention;

[0048] Figure 3 is a schematic diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0049] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0050] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in sequences other than those illustrated or described herein.

[0051] It should be understood that in various embodiments of the present invention, the size of the sequence number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0052] It should be understood that in the present invention, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products or apparatuses.

[0053] It should be understood that in the present invention, "plurality" refers to two or more than two. "And / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "Contains A, B and C", "Contains A, B, C" means that A, B, and C are all included, "Contains A, B or C" means that one of A, B, and C is included, and "Contains A, B and / or C" means that any one, any two, or any three of A, B, and C are included.

[0054] It should be understood that in the present invention, "B corresponding to A", "B corresponding to A", "A corresponds to B" or "B corresponds to A" means that B is associated with A and B can be determined based on A. Determining B based on A does not mean determining B based only on A, but B can also be determined based on A and / or other information. A and B match when the similarity between A and B is greater than or equal to a preset threshold.

[0055] Depending on the context, "if" as used herein may be interpreted as "when" or "when" or "in response to determining" or "in response to detecting."

[0056] The technical solution of the present invention is described in detail with specific embodiments below. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0057] In order to make the purpose, technical solutions and advantages of the present invention more clear, specific embodiments will be described below in conjunction with the accompanying drawings.

[0058] In one embodiment, Figure 1 As shown, a relocation method is provided, comprising the following steps:

[0059] Step S101: Obtain the semantic matrix and pose matrix corresponding to each frame image.

[0060] In order to obtain the semantic matrix corresponding to each frame image, it is necessary to first obtain the number of targets, target category labels and two-dimensional bounding box positions in each frame image, and then generate the semantic matrix of each frame image based on the number of targets, target category labels and two-dimensional bounding box positions in each frame image.

[0061] Specifically, a target detection algorithm (such as DETR) can be used to generate a semantic matrix The generated semantic matrix SM n as follows:

[0062]

[0063] Among them, b is the number of objects in each frame image, and the semantic matrix SM n The first column represents the target category label and the semantic matrix SM n The last four columns of represent the 2D bounding box locations.

[0064] To obtain the pose matrix corresponding to each frame of image, it is necessary to first obtain the rotation matrix, scaling factor and displacement of the camera in the world coordinate system, and then generate the pose matrix of each frame of image from the rotation matrix, scaling factor and displacement of the camera in the world coordinate system.

[0065] Specifically, the pose matrix The details are as follows:

[0066]

[0067] in, is the rotation matrix, is the scaling factor, is the displacement of the camera in the world coordinate system.

[0068] Step S102: Generate a pose semantic descriptor based on the semantic matrix and the pose matrix.

[0069] To generate a pose semantic descriptor based on the semantic matrix and pose matrix, it is necessary to first obtain the timestamp of each frame image, and then generate a pose semantic descriptor based on the semantic matrix and pose matrix corresponding to each frame image, as well as the timestamp of each frame image.

[0070] Among them, based on the semantic matrix and pose matrix corresponding to each frame image, as well as the timestamp of each frame image, a pose semantic descriptor is generated, including: extracting the first column data of the semantic matrix corresponding to each frame image, and forming a target category set from the first column data of the semantic matrix; extracting the last four columns data of the semantic matrix corresponding to each frame image, and forming a two-dimensional bounding box set from the last four columns data of the semantic matrix; generating a pose semantic descriptor from the timestamp of each frame image, the target category set, the two-dimensional bounding box set and the pose matrix.

[0071] Specifically, the pose semantic descriptor PSD n as follows:

[0072]

[0073] Where n represents the timestamp, represents the target category set, B n represents a set of two-dimensional bounding boxes, T wn Represents the pose matrix.

[0074] Step S103: Determine the target key frame image based on the pose semantic descriptor and preset constraints.

[0075] Among them, the preset constraints include preset posture constraints, preset category constraints and preset border constraints that are executed successively.

[0076] To determine the target key frame image based on the pose semantic descriptor and the preset constraints, it is necessary to first determine the first key frame image based on the pose semantic descriptor and the preset pose constraints, then determine the second key frame image based on the pose semantic descriptor, the preset category constraints and the first key frame image, and then determine the target key frame image based on the pose semantic descriptor, the preset border constraints and the second key frame image.

[0077] Among them, based on the pose semantic descriptor and the preset pose constraint conditions, determining the first key frame image includes: extracting the pose matrix in the pose semantic descriptor, and calculating the geometric similarity between the historical key frame image and the current key frame image based on the pose matrix; and selecting the key frame image whose geometric similarity satisfies the preset pose constraint conditions from the current key frame image as the first key frame image.

[0078] Specifically, the geometric similarity between the historical key frame image and the current key frame image is calculated first, and the calculation formula is as follows:

[0079] The camera pose is determined by T n,l Indicates that for two similar pose matrices T wn and T wl , define geometric similarity as the Frobenius normal form:

[0080] δT n,l =||T wn -T wl || F

[0081] Among them, ||·|| F represents the Frobenius norm, which is used to calculate the square root of the sum of squares of the elements of the matrix. δ represents a random number and can be set according to the specific situation.

[0082] The first key frame image to be screened must meet the following conditions:

[0083] ∈≤δT n,l ≤δT Th

[0084] Among them, δT Th is a predefined threshold, ∈ is used to take into account numerical errors, ∈=10 -4 , T n,l The camera pose is represented by a matrix, and δ represents a random number, which can be set according to the specific situation.

[0085] Wherein, based on the pose semantic descriptor, the preset category constraint and the first key frame image, determining the second key frame image includes: extracting the target category set in the pose semantic descriptor, and calculating the category label similarity between the historical key frame image and the current key frame image based on the target category set; and selecting the key frame image whose category label similarity satisfies the preset category constraint from the first key frame image as the second key frame image.

[0086] Specifically, the historical key frame images {C n} and the current key frame image {C j}C n,j , the formula is as follows:

[0087] δC n,j =|||c n ||2-||c j ||2|

[0088] Among them, ||·||2 is the L2 norm, which is used to measure the similarity of the category set, and δ represents a random number, which can be set according to the specific situation.

[0089] The second key frame image to be screened must meet the following conditions:

[0090] δC n,j ≤δC Th +0.16C Th

[0091] Among them, δCTh is an adaptive threshold, determined by the minimum category difference of all candidates, that is: δC Th =min({δC n,j}), C n,j Represents the historical key frame image {C n} and the current key frame image {C j}, δ represents a random number and can be set according to the specific situation.

[0092] Wherein, based on the pose semantic descriptor, the preset bounding box constraint and the second key frame image, the target key frame image is determined, including: extracting a two-dimensional bounding box set in the pose semantic descriptor, and calculating the overlap of the bounding boxes between the historical key frame image and the current key frame image based on the two-dimensional bounding box set; and selecting a key frame image whose overlap satisfies the preset bounding box constraint from the second key frame image as the target key frame image.

[0093] Specifically, the overlap of the bounding box between the historical key frame image and the current key frame image can be calculated by the intersection over union (IoU), and the formula is as follows:

[0094]

[0095] Among them, Area of ​​Overlap represents the overlapping area of ​​the bounding boxes between the historical key frame image and the current key frame image, and Area of ​​Union represents the union area of ​​the bounding boxes between the historical key frame image and the current key frame image.

[0096] It should be noted that only when the IoU is greater than the predefined threshold δIoU, the target is considered to overlap in two key frames and the counter v is increased. j .

[0097] The target key frame images to be screened must meet the following conditions:

[0098] v j ≥ b n

[0099] Among them, b n is the number of objects in the historical key frame images.

[0100] Step S104: performing relocation based on the target key frame image and the historical key frame images.

[0101] Once the target keyframe image is obtained, relocalization can be performed by solving the 3D to 2D PnP problem to restore the global pose in the active map.

[0102] An embodiment of the present invention provides a repositioning method, including: first obtaining a semantic matrix and a pose matrix corresponding to each frame image, then generating a pose semantic descriptor based on the semantic matrix and the pose matrix, and then determining a target key frame image based on the pose semantic descriptor and preset constraints, wherein the preset constraints include preset pose constraints, preset category constraints, and preset border constraints that are executed in sequence, and finally performing repositioning based on the target key frame image and the historical key frame image. The present application combines the semantic matrix and the pose matrix to generate a pose semantic descriptor, thereby screening the target key frame image through the pose semantic descriptor and the triple preset constraints to obtain an accurate target key frame image, thereby improving the accuracy of repositioning.

[0103] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0104] The following is an embodiment of the device of the present invention. For details not described in detail therein, reference may be made to the corresponding method embodiment described above.

[0105] Figure 2 The structure diagram of a relocation device provided by an embodiment of the present invention is shown. For the convenience of description, only the part related to the embodiment of the present invention is shown. The relocation device includes an acquisition module 201, a descriptor generation module 202, a target key frame determination module 203 and a relocation module 204, which are as follows:

[0106] An acquisition module 201 is used to acquire a semantic matrix and a pose matrix corresponding to each frame image;

[0107] A descriptor generation module 202, for generating a posture semantic descriptor based on a semantic matrix and a posture matrix;

[0108] A target key frame determination module 203 is used to determine a target key frame image based on a pose semantic descriptor and preset constraints, wherein the preset constraints include a preset pose constraint, a preset category constraint, and a preset border constraint that are executed in sequence;

[0109] The repositioning module 204 is used to perform repositioning based on the target key frame image and the historical key frame images.

[0110] In a possible implementation, the acquisition module 201 is further used to acquire the number of objects, object category labels, and two-dimensional bounding box positions in each frame image;

[0111] The semantic matrix of each frame image is generated based on the number of objects in each frame image, the object category label and the two-dimensional bounding box position.

[0112] In a possible implementation, the acquisition module 201 is further used to acquire a rotation matrix, a scaling factor, and a displacement of the camera in a world coordinate system;

[0113] The pose matrix of each frame image is generated by the rotation matrix, scaling factor and the displacement of the camera in the world coordinate system.

[0114] In a possible implementation, the descriptor generation module 202 is further used to obtain a timestamp of each frame image;

[0115] Based on the semantic matrix and pose matrix corresponding to each frame image, as well as the timestamp of each frame image, a pose semantic descriptor is generated.

[0116] In a possible implementation, the descriptor generation module 202 is further used to extract the first column data of the semantic matrix corresponding to each frame image, and form a target category set from the first column data of the semantic matrix;

[0117] Extract the last four columns of data of the semantic matrix corresponding to each frame image, and form a two-dimensional bounding box set from the last four columns of data of the semantic matrix;

[0118] The pose semantic descriptor is generated by the timestamp of each frame image, the set of target categories, the set of two-dimensional bounding boxes and the pose matrix.

[0119] In a possible implementation, the target key frame determination module 203 is further used to determine the first key frame image based on the pose semantic descriptor and the preset pose constraint condition;

[0120] Determine a second key frame image based on the pose semantic descriptor, the preset category constraint and the first key frame image;

[0121] Based on the pose semantic descriptor, the preset bounding box constraint and the second key frame image, a target key frame image is determined.

[0122] In a possible implementation, the target key frame determination module 203 is further used to extract a pose matrix in the pose semantic descriptor, and calculate the geometric similarity between the historical key frame image and the current key frame image based on the pose matrix;

[0123] A key frame image whose geometric similarity satisfies a preset position constraint condition is selected from the current key frame image as the first key frame image.

[0124] In a possible implementation, the target key frame determination module 203 is further used to extract a target category set in the pose semantic descriptor, and calculate the category label similarity between the historical key frame image and the current key frame image based on the target category set;

[0125] A key frame image whose category label similarity satisfies a preset category constraint condition is selected from the first key frame image as the second key frame image.

[0126] In a possible implementation, the target key frame determination module 203 is further used to extract a two-dimensional bounding box set in the pose semantic descriptor, and calculate the overlap of the bounding boxes between the historical key frame image and the current key frame image based on the two-dimensional bounding box set;

[0127] A key frame image whose overlap degree satisfies a preset border constraint condition is selected from the second key frame image as a target key frame image.

[0128] The embodiment of the present invention provides a repositioning device, which is specifically used to: first obtain the semantic matrix and pose matrix corresponding to each frame image, then generate a pose semantic descriptor based on the semantic matrix and the pose matrix, and then determine the target key frame image based on the pose semantic descriptor and preset constraints, wherein the preset constraints include preset pose constraints, preset category constraints, and preset border constraints that are executed in sequence, and finally perform repositioning based on the target key frame image and the historical key frame image. The present application combines the semantic matrix and the pose matrix to generate a pose semantic descriptor, thereby screening the target key frame image through the pose semantic descriptor and the triple preset constraints, obtaining an accurate target key frame image, and thus improving the accuracy of repositioning.

[0129] Figure 3 is a schematic diagram of a terminal provided by an embodiment of the present invention. Figure 3 As shown, the terminal 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, the steps in the above-mentioned various relocation method embodiments are implemented, such as Figure 1 Alternatively, when the processor 301 executes the computer program 303, the functions of each module / unit in the above-mentioned embodiments of the relocation device are realized, for example Figure 2 Functionality of modules / units 201-204 is shown.

[0130] The present invention further provides a readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, it is used to implement a relocation method provided by the above various embodiments, including:

[0131] Get the semantic matrix and pose matrix corresponding to each frame image;

[0132] Generate a pose semantic descriptor based on the semantic matrix and the pose matrix;

[0133] Determine a target key frame image based on the pose semantic descriptor and preset constraints, wherein the preset constraints include preset pose constraints, preset category constraints, and preset border constraints that are executed in sequence;

[0134] Relocalization is performed based on the target keyframe image and the historical keyframe images.

[0135] In a possible implementation, obtaining a semantic matrix corresponding to each frame image includes:

[0136] Get the number of objects, object category labels, and two-dimensional bounding box locations in each frame image;

[0137] The semantic matrix of each frame image is generated based on the number of objects in each frame image, the object category label and the two-dimensional bounding box position.

[0138] In a possible implementation, obtaining the pose matrix corresponding to each frame image includes:

[0139] Get the rotation matrix, scaling factor, and displacement of the camera in the world coordinate system;

[0140] The pose matrix of each frame image is generated by the rotation matrix, scaling factor and the displacement of the camera in the world coordinate system.

[0141] In a possible implementation, a pose semantic descriptor is generated based on a semantic matrix and a pose matrix, including:

[0142] Get the timestamp of each frame image;

[0143] Based on the semantic matrix and pose matrix corresponding to each frame image, as well as the timestamp of each frame image, a pose semantic descriptor is generated.

[0144] In a possible implementation, a pose semantic descriptor is generated based on the semantic matrix and pose matrix corresponding to each frame image and the timestamp of each frame image, including:

[0145] Extract the first column data of the semantic matrix corresponding to each frame image, and form a target category set based on the first column data of the semantic matrix;

[0146] Extract the last four columns of data of the semantic matrix corresponding to each frame image, and form a two-dimensional bounding box set from the last four columns of data of the semantic matrix;

[0147] The pose semantic descriptor is generated by the timestamp of each frame image, the set of target categories, the set of two-dimensional bounding boxes and the pose matrix.

[0148] In a possible implementation, determining a target key frame image based on a pose semantic descriptor and preset constraints includes:

[0149] Determine a first key frame image based on the pose semantic descriptor and preset pose constraints;

[0150] Determine a second key frame image based on the pose semantic descriptor, the preset category constraint and the first key frame image;

[0151] Based on the pose semantic descriptor, the preset bounding box constraint and the second key frame image, a target key frame image is determined.

[0152] In a possible implementation, determining a first key frame image based on a pose semantic descriptor and a preset pose constraint condition includes:

[0153] Extract the pose matrix in the pose semantic descriptor, and calculate the geometric similarity between the historical key frame image and the current key frame image based on the pose matrix;

[0154] A key frame image whose geometric similarity satisfies a preset position constraint condition is selected from the current key frame image as the first key frame image.

[0155] In a possible implementation, determining the second key frame image based on the pose semantic descriptor, the preset category constraint condition and the first key frame image includes:

[0156] Extract the target category set in the pose semantic descriptor, and calculate the category label similarity between the historical key frame image and the current key frame image based on the target category set;

[0157] A key frame image whose category label similarity satisfies a preset category constraint condition is selected from the first key frame image as the second key frame image.

[0158] In a possible implementation, determining a target key frame image based on a pose semantic descriptor, a preset bounding box constraint, and a second key frame image includes:

[0159] Extracting a two-dimensional bounding box set from the pose semantic descriptor, and calculating the overlap of the bounding boxes between the historical key frame image and the current key frame image based on the two-dimensional bounding box set;

[0160] A key frame image whose overlap degree satisfies a preset border constraint condition is selected from the second key frame image as a target key frame image.

[0161] Among them, the readable storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The computer storage medium can be any available medium that can be accessed by a general or special-purpose computer. For example, a readable storage medium is coupled to a processor so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application-specific integrated circuit (ASIC). In addition, the ASIC can be located in a user device. Of course, the processor and the readable storage medium can also exist in a communication device as discrete components. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0162] The present invention also provides a program product, which includes an execution instruction, and the execution instruction is stored in a readable storage medium. At least one processor of a device can read the execution instruction from the readable storage medium, and at least one processor executes the execution instruction so that the device implements a relocation method provided by the above various embodiments, including:

[0163] Get the semantic matrix and pose matrix corresponding to each frame image;

[0164] Generate a pose semantic descriptor based on the semantic matrix and the pose matrix;

[0165] Determine a target key frame image based on the pose semantic descriptor and preset constraints, wherein the preset constraints include preset pose constraints, preset category constraints, and preset border constraints that are executed in sequence;

[0166] Relocalization is performed based on the target keyframe image and the historical keyframe images.

[0167] In a possible implementation, obtaining a semantic matrix corresponding to each frame image includes:

[0168] Get the number of objects, object category labels, and two-dimensional bounding box locations in each frame image;

[0169] The semantic matrix of each frame image is generated based on the number of objects in each frame image, the object category label and the two-dimensional bounding box position.

[0170] In a possible implementation, obtaining the pose matrix corresponding to each frame image includes:

[0171] Get the rotation matrix, scaling factor, and displacement of the camera in the world coordinate system;

[0172] The pose matrix of each frame image is generated by the rotation matrix, scaling factor and the displacement of the camera in the world coordinate system.

[0173] In a possible implementation, a pose semantic descriptor is generated based on a semantic matrix and a pose matrix, including:

[0174] Get the timestamp of each frame image;

[0175] Based on the semantic matrix and pose matrix corresponding to each frame image, as well as the timestamp of each frame image, a pose semantic descriptor is generated.

[0176] In a possible implementation, a pose semantic descriptor is generated based on the semantic matrix and pose matrix corresponding to each frame image and the timestamp of each frame image, including:

[0177] Extract the first column data of the semantic matrix corresponding to each frame image, and form a target category set based on the first column data of the semantic matrix;

[0178] Extract the last four columns of data of the semantic matrix corresponding to each frame image, and form a two-dimensional bounding box set from the last four columns of data of the semantic matrix;

[0179] The pose semantic descriptor is generated by the timestamp of each frame image, the set of target categories, the set of two-dimensional bounding boxes and the pose matrix.

[0180] In a possible implementation, determining a target key frame image based on a pose semantic descriptor and preset constraints includes:

[0181] Determine a first key frame image based on the pose semantic descriptor and preset pose constraints;

[0182] Determine a second key frame image based on the pose semantic descriptor, the preset category constraint and the first key frame image;

[0183] Based on the pose semantic descriptor, the preset bounding box constraint and the second key frame image, a target key frame image is determined.

[0184] In a possible implementation, determining a first key frame image based on a pose semantic descriptor and a preset pose constraint condition includes:

[0185] Extract the pose matrix in the pose semantic descriptor, and calculate the geometric similarity between the historical key frame image and the current key frame image based on the pose matrix;

[0186] A key frame image whose geometric similarity satisfies a preset position constraint condition is selected from the current key frame image as the first key frame image.

[0187] In a possible implementation, determining the second key frame image based on the pose semantic descriptor, the preset category constraint condition and the first key frame image includes:

[0188] Extract the target category set in the pose semantic descriptor, and calculate the category label similarity between the historical key frame image and the current key frame image based on the target category set;

[0189] A key frame image whose category label similarity satisfies a preset category constraint condition is selected from the first key frame image as the second key frame image.

[0190] In a possible implementation, determining a target key frame image based on a pose semantic descriptor, a preset bounding box constraint, and a second key frame image includes:

[0191] Extracting a two-dimensional bounding box set from the pose semantic descriptor, and calculating the overlap of the bounding boxes between the historical key frame image and the current key frame image based on the two-dimensional bounding box set;

[0192] A key frame image whose overlap degree satisfies a preset border constraint condition is selected from the second key frame image as a target key frame image.

[0193] In the embodiments of the above-mentioned devices, it should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. The general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc. The steps of the method disclosed in the present invention may be directly embodied as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0194] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A relocation method, characterized in that: include: Get the semantic matrix and pose matrix corresponding to each frame image; Generate a pose semantic descriptor based on the semantic matrix and the pose matrix; Determine a target key frame image based on the pose semantic descriptor and preset constraints, wherein the preset constraints include preset pose constraints, preset category constraints, and preset border constraints that are executed sequentially; Relocalization is performed based on the target key frame image and the historical key frame images.

2. A relocation method according to claim 1, characterized in that: The obtaining of the semantic matrix corresponding to each frame image includes: Obtaining the number of objects in each frame of image, object category labels, and two-dimensional bounding box positions; A semantic matrix of each frame of image is generated based on the number of objects in each frame of image, the object category label and the two-dimensional bounding box position.

3. A relocation method according to claim 1, characterized in that: The obtaining of the pose matrix corresponding to each frame of image includes: Get the rotation matrix, scaling factor, and displacement of the camera in the world coordinate system; The position matrix of each frame image is generated by the rotation matrix, the scaling factor and the displacement of the camera in the world coordinate system.

4. A relocation method according to claim 1, characterized in that: The step of generating a pose semantic descriptor based on the semantic matrix and the pose matrix comprises: Get the timestamp of each frame image; The pose semantic descriptor is generated based on the semantic matrix and pose matrix corresponding to each frame image and the timestamp of each frame image.

5. A relocation method according to claim 4, characterized in that: The generating the pose semantic descriptor based on the semantic matrix and pose matrix corresponding to each frame image and the timestamp of each frame image includes: Extracting the first column data of the semantic matrix corresponding to each frame image, and forming a target category set from the first column data of the semantic matrix; Extracting the last four columns of data of the semantic matrix corresponding to each frame image, and forming a two-dimensional bounding box set by the last four columns of data of the semantic matrix; The pose semantic descriptor is generated by the timestamps of the frames of images, the target category set, the two-dimensional bounding box set and the pose matrix.

6. A relocation method according to claim 1, characterized in that: The step of determining a target key frame image based on the pose semantic descriptor and preset constraints includes: Determining a first key frame image based on the pose semantic descriptor and the preset pose constraint condition; Determining a second key frame image based on the pose semantic descriptor, the preset category constraint and the first key frame image; The target key frame image is determined based on the pose semantic descriptor, the preset bounding box constraint and the second key frame image.

7. A relocation method according to claim 6, characterized in that: The determining of the first key frame image based on the posture semantic descriptor and the preset posture constraint condition includes: Extracting a pose matrix from the pose semantic descriptor, and calculating the geometric similarity between the historical key frame image and the current key frame image based on the pose matrix; A key frame image whose geometric similarity satisfies a preset position constraint condition is selected from the current key frame image as the first key frame image.

8. A relocation method according to claim 6, characterized in that: The determining of the second key frame image based on the pose semantic descriptor, the preset category constraint and the first key frame image includes: Extracting a target category set from the pose semantic descriptor, and calculating the category label similarity between the historical key frame image and the current key frame image based on the target category set; A key frame image whose category label similarity satisfies a preset category constraint condition is selected from the first key frame images as the second key frame image.

9. A relocation method according to claim 6, characterized in that: The step of determining the target key frame image based on the pose semantic descriptor, the preset bounding box constraint and the second key frame image includes: Extracting a two-dimensional bounding box set in the pose semantic descriptor, and calculating the degree of overlap of the bounding boxes between the historical key frame image and the current key frame image based on the two-dimensional bounding box set; A key frame image whose overlap degree satisfies a preset frame constraint condition is selected from the second key frame images as the target key frame image.

10. A repositioning device, characterized in that: include: An acquisition module is used to obtain the semantic matrix and posture matrix corresponding to each frame image; A descriptor generation module, used for generating a posture semantic descriptor based on the semantic matrix and the posture matrix; A target key frame determination module, used to determine a target key frame image based on the pose semantic descriptor and preset constraints, wherein the preset constraints include preset pose constraints, preset category constraints, and preset border constraints that are executed in sequence; A repositioning module is used to perform repositioning based on the target key frame image and the historical key frame images.