Video positioning method based on information entropy screening and feature point matching
By introducing information entropy screening and feature point matching in the video positioning method, combining deep learning and geometric constraint model, the problems of low accuracy and complexity of traditional indoor positioning methods are solved, and indoor positioning with centimeter-level accuracy is achieved.
Patent Information
- Application Number
- CN202510255788.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-20
AI Technical Summary
Traditional image-based indoor positioning methods have problems such as low accuracy, large delay, and complex operation, which are difficult to meet the needs of high-precision scenarios.
The video positioning method based on information entropy screening and matching feature points is adopted, feature vectors are extracted through deep learning networks, and keyframes are extracted in combination with grayscale information entropy and time interval limitations, and the geometric constraint model is used to achieve decimeter-level accuracy positioning.
It significantly improves positioning accuracy, reduces calculation complexity and delay, and is suitable for high-precision positioning scenarios such as smart home, industrial logistics, and AR/VR.
Smart Images

Figure CN120182374A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer and software technologies, and particularly to a video positioning method. Background Art
[0002] With the rapid development of information technology and the Internet of Things, indoor positioning technology, as an important part of modern positioning and navigation systems, has gradually become the core support technology in intelligent scenarios. It is widely used in fields such as smart home, personalized services, industrial logistics, retail analysis, healthcare, and public safety, and is of great significance for improving production efficiency, optimizing resource allocation, and enhancing user experience. Especially in smart home, indoor positioning technology can achieve efficient resource scheduling by real-time tracking the positions of devices, robots, and personnel; in the field of industrial logistics, it can optimize production processes and improve logistics efficiency; in the field of retail analysis, it helps merchants obtain customer behavior data for personalized marketing.
[0003] In addition, indoor positioning technology also has significant applications in the healthcare field. By helping to quickly locate medical devices and patients, it greatly saves emergency time and improves the efficiency of medical services. In the office environment, indoor positioning technology can be used for dynamic allocation of office workstations and parking spaces, avoiding equipment loss or idleness, and improving the utilization efficiency of office space. At the same time, indoor positioning technology can also provide precise navigation services for scenarios such as shopping malls and airports, bringing immersive experiences in combination with augmented reality (AR) and virtual reality (VR) technologies. In public safety and emergency response, it can monitor crowd flow, quickly locate trapped people, and ensure personnel safety.
[0004] However, although existing wireless signal positioning technologies have achieved certain results in some fields, the indoor positioning method based on images has gradually become a research hotspot due to its unique advantages. Compared with traditional wireless signal positioning methods, the indoor positioning technology based on images has high environmental adaptability and intuitive visualization effects, and can use image features in the environment for positioning, especially performing well in complex and dynamically changing scenarios. In addition, image-based methods usually can use existing devices (such as the cameras of smartphones) for positioning, with relatively low hardware costs and wide application ranges.
[0005] However, traditional image-based positioning methods also have some deficiencies. First, users need to manually take multiple images near the location and label landmarks, which not only increases the complexity of operations but also easily introduces human errors. Second, limited by the performance of image matching algorithms and interference factors such as environmental light changes and occlusions, the positioning accuracy of traditional image positioning methods usually can only reach the meter level or sub-meter level, making it difficult to meet the requirements of high-precision scenarios. Especially in application scenarios with extremely high requirements for positioning accuracy, such as industrial robot navigation, augmented reality (AR), and virtual reality (VR), the meter-level accuracy obviously cannot meet the centimeter-level accuracy requirements, resulting in the system accuracy failing to achieve the expected effect.
[0006] Therefore, there is an urgent need to study an efficient, convenient, and centimeter-level accurate image positioning method to overcome the problems of low accuracy, large time delay, and complex operations existing in traditional positioning technologies. Summary of the Invention
[0007] The purpose of the present invention is to solve the problems of low accuracy, large time delay, and complex operations existing in traditional positioning technologies, and propose a video positioning method based on information entropy screening and feature point matching.
[0008] The specific process of the video positioning method based on information entropy screening and feature point matching is as follows:
[0009] Step 1: Construct an indoor image sample set with unlabeled categories and labeled location information;
[0010] Input the indoor image sample set with unlabeled categories and labeled location information into the backbone network of the deep learning network model. The backbone network of the deep learning network model outputs feature vectors, and obtain the feature vectors of each image in the indoor image sample set;
[0011] Step 2: Collect video data and synchronously record the rotation angle data of the camera;
[0012] Step 3: Extract key frame images from the video data based on gray information entropy and combined with time interval limitations;
[0013] Step 4:
[0014] Input the key frame images extracted in Step 3 into the backbone network of the deep learning network model. The backbone network of the deep learning network model outputs the feature vectors of the key frame images;
[0015] Calculate the similarity between the feature vector of the i-th key frame image and the feature vectors of each image in the indoor image sample set obtained in Step 1;
[0016] Select the three key frames with the highest similarity to the feature vector of the i-th key frame image and the position information corresponding to the three key frames from the feature vectors of each image in the indoor image sample set obtained in Step 1 as the matching result of the i-th key frame image;
[0017] Until the matching results of all key frame images are obtained;
[0018] Step 5: Locate the camera position in the map based on the matching results of the key frame images.
[0019] Preferably, the deep learning network model includes: a backbone network and a fully connected layer;
[0020] The backbone network includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a Maxout activation function layer, a pooling layer, a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer, a thirteenth convolutional layer, a fourteenth convolutional layer, a fifteenth convolutional layer, a sixteenth convolutional layer, a seventeenth convolutional layer, an eighteenth convolutional layer, a nineteenth convolutional layer, a twentieth convolutional layer, a twenty-first convolutional layer, a twenty-second convolutional layer, a first gating mechanism, a second gating mechanism, a third gating mechanism, a fourth gating mechanism, a fifth gating mechanism, a sixth gating mechanism, a seventh gating mechanism, an eighth gating mechanism;
[0021] Each of the first gating mechanism, the second gating mechanism, the third gating mechanism, the fourth gating mechanism, the fifth gating mechanism, the sixth gating mechanism, the seventh gating mechanism, and the eighth gating mechanism sequentially includes a 1×1 convolutional layer and a Sigmoid activation function layer;
[0022] The convolution kernel size of the first convolutional layer is 3×1, and the number of channels is 64;
[0023] The convolution kernel size of the second convolutional layer is 5×1, and the number of channels is 64;
[0024] The convolution kernel size of the third convolutional layer is 7×1, and the number of channels is 64;
[0025] The convolution kernel size of the fourth convolutional layer is 1×3, and the number of channels is 64;
[0026] The convolution kernel size of the fifth convolutional layer is 1×5, and the number of channels is 64;
[0027] The convolution kernel size of the sixth convolutional layer is 1×7, and the number of channels is 64;
[0028] The convolution kernel size of the seventh convolutional layer is 3×3, the number of channels is 64, and the stride is 2;
[0029] The convolution kernel size of the eighth convolutional layer is 3×3, the number of channels is 64, and the stride is 2;
[0030] The convolutional kernel size of the ninth convolutional layer is 3×3, the number of channels is 64, and the stride is 2;
[0031] The convolutional kernel size of the tenth convolutional layer is 3×3, the number of channels is 64, and the stride is 2;
[0032] The convolutional kernel size of the eleventh convolutional layer is 3×3, the number of channels is 128, and the stride is 2;
[0033] The convolutional kernel size of the twelfth convolutional layer is 3×3, the number of channels is 128, and the stride is 2;
[0034] The convolutional kernel size of the thirteenth convolutional layer is 3×3, the number of channels is 128, and the stride is 2;
[0035] The convolutional kernel size of the fourteenth convolutional layer is 3×3, the number of channels is 128, and the stride is 2;
[0036] The convolutional kernel size of the fifteenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2;
[0037] The convolutional kernel size of the sixteenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2;
[0038] The convolutional kernel size of the seventeenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2;
[0039] The convolutional kernel size of the eighteenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2;
[0040] The convolutional kernel size of the nineteenth convolutional layer is 3×3, the number of channels is 512, and the stride is 2;
[0041] The convolutional kernel size of the twentieth convolutional layer is 3×3, the number of channels is 512, and the stride is 2;
[0042] The convolutional kernel size of the twenty - first convolutional layer is 3×3, the number of channels is 512, and the stride is 2;
[0043] The convolutional kernel size of the twenty - second convolutional layer is 3×3, the number of channels is 512, and the stride is 2.
[0044] Preferably, the working process of the deep learning network model is as follows:
[0045] Images are respectively input into the first convolutional layer, the second convolutional layer, and the third convolutional layer of the deep learning network model. The first convolutional layer outputs feature α, the second convolutional layer outputs feature α′, and the third convolutional layer outputs feature α″;
[0046] The feature α output by the first convolutional layer is input into the fourth convolutional layer, and the fourth convolutional layer outputs feature α″′;
[0047] The output feature α′ of the second convolutional layer is input into the fifth convolutional layer, and the fifth convolutional layer outputs a feature
[0048] The output feature α″ of the third convolutional layer is input into the sixth convolutional layer, and the sixth convolutional layer outputs a feature
[0049] The output feature α′″ of the fourth convolutional layer and the output feature of the fifth convolutional layer The output feature of the sixth convolutional layer are input into the Maxout activation function layer, and the Maxout activation function layer outputs a feature α ;
[0050] The output feature of the Maxout activation function layer α is input into the pooling layer, and the pooling layer outputs a feature β;
[0051] The output feature β of the pooling layer is sequentially input into the seventh convolutional layer and the eighth convolutional layer, and the eighth convolutional layer outputs a feature β′;
[0052] The output feature β of the pooling layer is input into the first gating mechanism, and the first gating mechanism outputs a feature β″;
[0053] The output feature β′ of the eighth convolutional layer and the output feature β″ of the first gating mechanism are sequentially input into the ninth convolutional layer and the tenth convolutional layer, and the tenth convolutional layer outputs a feature β″′;
[0054] The output feature β′ of the eighth convolutional layer and the output feature β″ of the first gating mechanism are input into the second gating mechanism, and the second gating mechanism outputs a feature
[0055] The output feature β″′ of the tenth convolutional layer and the output feature of the second gating mechanism are sequentially input into the eleventh convolutional layer and the twelfth convolutional layer, and the twelfth convolutional layer outputs a feature
[0056] The output feature β″′ of the tenth convolutional layer and the output feature of the second gating mechanism are input into the third gating mechanism, and the third gating mechanism outputs a feature β ;
[0057] The output feature of the twelfth convolutional layer and the output feature of the third gating mechanism β are sequentially input into the thirteenth convolutional layer and the fourteenth convolutional layer, and the fourteenth convolutional layer outputs a feature γ;
[0058] The output feature of the twelfth convolutional layer and the output feature of the third gating mechanism β are input into the fourth gating mechanism, and the fourth gating mechanism outputs a feature γ′;
[0059] The output feature γ of the fourteenth convolutional layer and the output feature γ′ of the fourth gating mechanism are successively input into the fifteenth convolutional layer and the sixteenth convolutional layer, and the sixteenth convolutional layer outputs the feature γ″;
[0060] The output feature γ of the fourteenth convolutional layer and the output feature γ′ of the fourth gating mechanism are input into the fifth gating mechanism, and the fifth gating mechanism outputs the feature γ‴;
[0061] The output feature γ″ of the sixteenth convolutional layer and the output feature γ‴ of the fifth gating mechanism are successively input into the seventeenth convolutional layer and the eighteenth convolutional layer, and the eighteenth convolutional layer outputs the feature
[0062] The output feature γ″ of the sixteenth convolutional layer and the output feature γ‴ of the fifth gating mechanism are input into the sixth gating mechanism, and the sixth gating mechanism outputs the feature
[0063] The output feature of the eighteenth convolutional layer and the output feature of the sixth gating mechanism γ ;
[0064] The output feature of the eighteenth convolutional layer and the output feature of the sixth gating mechanism are input into the seventh gating mechanism, and the seventh gating mechanism outputs the feature δ;
[0065] The output feature γ of the twentieth convolutional layer and the output feature δ of the seventh gating mechanism are successively input into the twenty - first convolutional layer and the twenty - second convolutional layer, and the twenty - second convolutional layer outputs the feature δ;
[0066] The output feature γ of the twentieth convolutional layer and the output feature δ of the seventh gating mechanism are input into the eighth gating mechanism, and the eighth gating mechanism outputs the feature δ″;
[0067] The feature vector after splicing the output feature δ′ of the twenty - second convolutional layer and the output feature δ″ of the eighth gating mechanism is the output feature vector of the backbone network in the deep learning network model;
[0068] The output feature vector of the backbone network is input into the fully - connected layer, and the fully - connected layer outputs the classification result.
[0069] Preferably, the deep learning network model is a pre - trained deep learning network model;
[0070] The pre - training process of the deep learning network model is as follows:
[0071] Input the image set with labeled category and location information into the backbone network of the deep learning network model. The backbone network of the deep learning network model outputs feature vectors, and input the feature vectors of each image into the fully connected layer for classification;
[0072] Until the entire image set with labeled category and location information is input into the backbone network, a pre-trained deep learning network model is obtained.
[0073] Preferably, in the video data acquisition in step two, the rotation angle data of the camera is recorded synchronously; the specific process is as follows:
[0074] Use a gyroscope sensor to record the rotation angle data of the camera when shooting the video.
[0075] Preferably, in step three, key frame images in the video data are extracted based on gray information entropy and combined with time interval limitation; the specific process is as follows:
[0076] Calculate the gray information entropy of each frame in the video data, sort all frames in descending order of entropy value to obtain the sorted frame numbers;
[0077] Set the minimum frame interval threshold;
[0078] Delete the frame numbers with frame intervals less than the minimum frame interval threshold among the sorted frame numbers, retain the frame numbers greater than or equal to the minimum frame interval threshold, and obtain the key frame images in the video data.
[0079] Preferably, the process of deleting the frame numbers with frame intervals less than the minimum frame interval threshold among the sorted frame numbers, retaining the frame numbers greater than or equal to the minimum frame interval threshold, and obtaining the key frame images in the video data; the specific process is as follows:
[0080] The frame numbers after all frames are sorted in descending order of entropy value are A, B, C, D, E, F, G;
[0081] If the frame interval between frame number A and frame number B is less than the minimum frame interval, discard frame number B, and continue to compare frame number A with frame number C until all frame numbers are traversed;
[0082] If the frame interval between frame number A and frame number B is greater than or equal to the minimum frame interval, retain frame number A and frame number B, and continue to compare frame number B with frame number C until all frame numbers are traversed;
[0083] Obtain the key frame images in the video data.
[0084] Preferably, in step four, the similarity is the cosine value of the included angle between feature vectors.
[0085] Preferably, based on the key-frame matching results, the camera position is located in the map; the specific process is as follows:
[0086] Mark the positions of the three key frames in the matching result as point A(x1, y1), B(x2, y2), and C(x3, y3) respectively;
[0087] The included angle θ1 between the target point P(x, y) and the reference points A(x1, y1) and B(x2, y2) is calculated through the vector dot product formula:
[0088]
[0089] The included angle θ2 between the target point P(x, y) and the reference points B(x2, y2) and C(x3, y3) is calculated through the vector dot product formula:
[0090]
[0091] Wherein,
[0092] Vector PA represents the connection line from the target point P(x, y) to the reference point A(x1, y1);
[0093] Vector PB represents the connection line from the target point P(x, y) to the reference point B(x2, y2);
[0094] Vector PC represents the connection line from the target point P(x, y) to the reference point C(x3, y3);
[0095] The target point P(x, y) is the camera position;
[0096] PA·PB represents the dot product of vector PA and vector PB;
[0097] PC·PB represents the dot product of vector PC and vector PB;
[0098] ||PA|| represents the magnitude of vector PA;
[0099] ||PB|| represents the magnitude of vector PB;
[0100] ||PC|| represents the magnitude of vector PC;
[0101] According to the included angle constraint, a non-linear equation system is obtained:
[0102]
[0103] Wherein,
[0104] f1(x, y) represents the constraint condition of the included angle between vector PA and PB;
[0105] f2(x,y) represents the constraint condition of the included angle between vectors PB and PC.
[0106] The Gauss-Newton method is used to iteratively solve the non-linear equations to obtain the position P(x,y) of the target point.
[0107] Preferably, the dot product of vector PA and vector PB is expressed as: PA·PB = (x1 - x)(x2 - x) + (y1 - y)(y2 - y);
[0108] The dot product of vector PC and vector PB is expressed as: PC·PB = (x3 - x)(x2 - x) + (y3 - y)(y2 - y);
[0109] The magnitude ||PA|| of vector PA is expressed as:
[0110] The magnitude ||PB|| of vector PB is expressed as:
[0111] The magnitude ||PC|| of vector PC is expressed as:
[0112] The beneficial effects of the present invention are:
[0113] The present invention proposes a video positioning method based on information entropy screening and feature point matching. By combining deep learning, image processing, and geometric constraint technologies, centimeter-level indoor positioning is achieved. This method uses gray information entropy to screen key frames, extracts feature vectors through a deep neural network for matching, combines the camera angle information recorded by the gyroscope, and uses a geometric constraint model to achieve decimeter-level positioning accuracy. This method significantly reduces the complexity and latency of traditional image positioning and is applicable to high-precision positioning scenarios such as smart homes, industrial logistics, and AR / VR.
[0114] The present invention proposes a video positioning method combining information entropy, image matching, and geometric constraints, which has significant beneficial effects. Through the following technical means, the present invention effectively solves multiple challenges in traditional video positioning technologies and makes breakthrough progress in multiple aspects:
[0115] The present invention can significantly solve problems such as high complexity and large latency existing in traditional image positioning technologies;
[0116] By using an efficient deep neural network for image matching, the present invention reduces the inference complexity while ensuring the matching accuracy; by designing neural networks of different sizes, the requirements of different computing power devices can be adapted, further improving the adaptability and universality of the system.
[0117] The present invention proposes to store only the feature vectors of the data in the image library, significantly reducing the storage space requirement (limited to the MB level), accelerating the matching process, and improving the real-time performance and stability of the system.
[0118] The present invention realizes the decimeter-level positioning of the camera position by combining geometric constraints, further improving the positioning accuracy and fineness, and meeting the requirements of high-precision indoor positioning.
[0119] In summary, the method of the present invention has significant advantages in improving positioning accuracy, reducing computational burden, reducing communication latency, and optimizing device adaptation, and is widely applicable to various video positioning applications. Although there are challenges in device adaptation, environmental adaptability, and computational resource optimization, the technical means provided by the present invention have achieved good results in the experimental environment, have strong practical value and broad application prospects, especially in indoor navigation, intelligent monitoring, AR / VR applications and other fields have important application potential. Brief Description of the Drawings
[0120] Figure 1 is a flowchart of the implementation of the present invention;
[0121] Figure 2 is a structural diagram of the backbone network in the deep learning network model of the present invention. Detailed Embodiments
[0122] Detailed Embodiment 1: The specific process of the video positioning method based on information entropy screening and feature point matching in this embodiment is as follows:
[0123] Step 1: Construct an indoor image sample set with unlabeled categories and labeled position information;
[0124] Input the indoor image sample set with unlabeled categories and labeled position information into the backbone network of the deep learning network model. The backbone network of the deep learning network model outputs feature vectors, and the feature vectors of each image in the indoor image sample set are obtained;
[0125] Step 2: Collect video data and synchronously record the rotation angle data of the camera;
[0126] Step 3: Extract key frame images from the video data based on gray information entropy and combined with time interval constraints;
[0127] Step 4:
[0128] Input the key frame images extracted in Step 3 into the backbone network of the deep learning network model. The backbone network of the deep learning network model outputs the feature vectors of the key frame images;
[0129] Calculate the similarity between the feature vector of the i-th key frame image and the feature vector of each image in the indoor image sample set obtained in Step 1;
[0130] Select three key frames with the highest similarity to the feature vector of the i-th key frame image and the position information corresponding to the three key frames from the feature vectors of each image in the indoor image sample set obtained in Step 1 as the matching result of the i-th key frame image;
[0131] Until the matching results of all key frame images are obtained;
[0132] Step 5: Locate the camera position in the map based on the matching results of the key frame images.
[0133] Specific Embodiment 2: The difference between this embodiment and Specific Embodiment 1 is that the deep learning network model includes: a backbone network and a fully connected layer;
[0134] The backbone network includes: a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a Maxout activation function layer, a pooling layer, a seventh convolutional layer, an eighth convolutional layer, a ninth convolutional layer, a tenth convolutional layer, an eleventh convolutional layer, a twelfth convolutional layer, a thirteenth convolutional layer, a fourteenth convolutional layer, a fifteenth convolutional layer, a sixteenth convolutional layer, a seventeenth convolutional layer, an eighteenth convolutional layer, a nineteenth convolutional layer, a twentieth convolutional layer, a twenty-first convolutional layer, a twenty-second convolutional layer, a first gating mechanism, a second gating mechanism, a third gating mechanism, a fourth gating mechanism, a fifth gating mechanism, a sixth gating mechanism, a seventh gating mechanism, an eighth gating mechanism;
[0135] Each of the first gating mechanism, the second gating mechanism, the third gating mechanism, the fourth gating mechanism, the fifth gating mechanism, the sixth gating mechanism, the seventh gating mechanism, and the eighth gating mechanism sequentially includes a 1×1 convolutional layer and a Sigmoid activation function layer;
[0136] The convolutional kernel size of the first convolutional layer is 3×1 (capturing horizontal features), and the number of channels is 64;
[0137] The convolutional kernel size of the second convolutional layer is 5×1 (capturing horizontal features), and the number of channels is 64;
[0138] The convolutional kernel size of the third convolutional layer is 7×1 (capturing horizontal features), and the number of channels is 64;
[0139] The convolutional kernel size of the fourth convolutional layer is 1×3 (capturing vertical features), and the number of channels is 64;
[0140] The convolutional kernel size of the fifth convolutional layer is 1×5 (capturing vertical features), and the number of channels is 64;
[0141] The convolutional kernel size of the sixth convolutional layer is 1×7 (capturing vertical features), and the number of channels is 64;
[0142] The convolution kernel size of the seventh convolutional layer is 3×3, the number of channels is 64, and the stride is 2;
[0143] The convolution kernel size of the eighth convolutional layer is 3×3, the number of channels is 64, and the stride is 2;
[0144] The convolution kernel size of the ninth convolutional layer is 3×3, the number of channels is 64, and the stride is 2;
[0145] The convolution kernel size of the tenth convolutional layer is 3×3, the number of channels is 64, and the stride is 2;
[0146] The convolution kernel size of the eleventh convolutional layer is 3×3, the number of channels is 128, and the stride is 2;
[0147] The convolution kernel size of the twelfth convolutional layer is 3×3, the number of channels is 128, and the stride is 2;
[0148] The convolution kernel size of the thirteenth convolutional layer is 3×3, the number of channels is 128, and the stride is 2;
[0149] The convolution kernel size of the fourteenth convolutional layer is 3×3, the number of channels is 128, and the stride is 2;
[0150] The convolution kernel size of the fifteenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2;
[0151] The convolution kernel size of the sixteenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2;
[0152] The convolution kernel size of the seventeenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2;
[0153] The convolution kernel size of the eighteenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2;
[0154] The convolution kernel size of the nineteenth convolutional layer is 3×3, the number of channels is 512, and the stride is 2;
[0155] The convolution kernel size of the twentieth convolutional layer is 3×3, the number of channels is 512, and the stride is 2;
[0156] The convolution kernel size of the twenty - first convolutional layer is 3×3, the number of channels is 512, and the stride is 2;
[0157] The convolution kernel size of the twenty - second convolutional layer is 3×3, the number of channels is 512, and the stride is 2.
[0158] Other steps and parameters are the same as those in the specific implementation method one.
[0159] Specific implementation method three: The difference between this implementation method and the first or second specific implementation method is that the working process of the deep learning network model is as follows:
[0160] The images are respectively input into the first convolutional layer, the second convolutional layer, and the third convolutional layer of the deep learning network model. The first convolutional layer outputs feature α, the second convolutional layer outputs feature α′, and the third convolutional layer outputs feature α″;
[0161] The feature α output by the first convolutional layer is input into the fourth convolutional layer, and the fourth convolutional layer outputs feature α″′;
[0162] The feature α′ output by the second convolutional layer is input into the fifth convolutional layer, and the fifth convolutional layer outputs feature
[0163] The feature α″ output by the third convolutional layer is input into the sixth convolutional layer, and the sixth convolutional layer outputs feature
[0164] The feature α″′ output by the fourth convolutional layer, the feature output by the sixth convolutional layer are input into the Maxout activation function layer, and the Maxout activation function layer outputs feature α ;
[0165] The feature output by the Maxout activation function layer α is input into the pooling layer, and the pooling layer outputs feature β;
[0166] The feature β output by the pooling layer is successively input into the seventh convolutional layer and the eighth convolutional layer, and the eighth convolutional layer outputs feature β′;
[0167] The feature β output by the pooling layer is input into the first gating mechanism, and the first gating mechanism outputs feature β″;
[0168] The feature β′ output by the eighth convolutional layer and the feature β″ output by the first gating mechanism are successively input into the ninth convolutional layer and the tenth convolutional layer, and the tenth convolutional layer outputs feature β′″;
[0169] The feature β′ output by the eighth convolutional layer and the feature β″ output by the first gating mechanism are input into the second gating mechanism, and the second gating mechanism outputs feature
[0170] The feature β′″ output by the tenth convolutional layer and the feature output by the second gating mechanism are successively input into the eleventh convolutional layer and the twelfth convolutional layer, and the twelfth convolutional layer outputs feature
[0171] The feature β′″ output by the tenth convolutional layer and the feature output by the second gating mechanism are input into the third gating mechanism, and the third gating mechanism outputs feature β ;
[0172] The feature output by the twelfth convolutional layer and the output feature of the third gating mechanism β are sequentially input into the thirteenth convolutional layer and the fourteenth convolutional layer, and the output feature γ of the fourteenth convolutional layer is obtained;
[0173] The output feature of the twelfth convolutional layer β and the output feature of the third gating mechanism are input into the fourth gating mechanism, and the output feature γ′ of the fourth gating mechanism is obtained;
[0174] The output feature γ of the fourteenth convolutional layer and the output feature γ′ of the fourth gating mechanism are sequentially input into the fifteenth convolutional layer and the sixteenth convolutional layer, and the output feature γ″ of the sixteenth convolutional layer is obtained;
[0175] The output feature γ of the fourteenth convolutional layer and the output feature γ′ of the fourth gating mechanism are input into the fifth gating mechanism, and the output feature γ″′ of the fifth gating mechanism is obtained;
[0176] The output feature γ″ of the sixteenth convolutional layer and the output feature γ″′ of the fifth gating mechanism are sequentially input into the seventeenth convolutional layer and the eighteenth convolutional layer, and the output feature
[0177] The output feature γ″ of the sixteenth convolutional layer and the output feature γ″′ of the fifth gating mechanism are input into the sixth gating mechanism, and the output feature
[0178] The output feature of the eighteenth convolutional layer and the output feature of the sixth gating mechanism are sequentially input into the nineteenth convolutional layer and the twentieth convolutional layer, and the output feature γ ;
[0179] The output feature of the eighteenth convolutional layer and the output feature of the sixth gating mechanism are input into the seventh gating mechanism, and the output feature δ is obtained;
[0180] The output feature γ of the twentieth convolutional layer and the output feature δ of the seventh gating mechanism are sequentially input into the twenty-first convolutional layer and the twenty-second convolutional layer, and the output feature δ′ of the twenty-second convolutional layer is obtained;
[0181] The output feature γ of the twentieth convolutional layer and the output feature δ of the seventh gating mechanism are input into the eighth gating mechanism, and the output feature δ″ is obtained;
[0182] The feature vector obtained by concatenating the output feature δ′ of the twenty-second convolutional layer and the output feature δ″ of the eighth gating mechanism is the output feature vector of the backbone network in the deep learning network model;
[0183] The output feature vector of the backbone network is input into the fully connected layer, and the fully connected layer outputs the classification result.
[0184] The present invention proposes an improved deep learning network method based on residuals, aiming to optimize the performance of traditional residual networks in multi-scale feature extraction by introducing a multi-convolution kernel module and a gating mechanism. Specifically, in the process of processing input data, this method adopts a multi-convolution kernel module, which effectively extracts features of the input data at different scales by applying multiple sizes of convolution kernels in parallel, thereby enhancing the model's perception and expression ability of different sizes and multi-scale features.
[0185] The multi-convolution kernel module includes multiple convolution kernels with different sizes. These convolution kernels work in parallel at the same level to extract features of the input image or data from multiple angles. The multiple feature maps obtained through parallel convolution operations will be fused and formed into a unified feature representation through concatenation operations. In this way, local information of multiple scales can be captured simultaneously in the same layer, thereby improving the network's expression ability for multi-scale features. Compared with the traditional single-convolution kernel structure, the multi-convolution kernel module can capture richer and more diverse feature information and improve the model's processing ability for complex data. At the output of this module, the Maxout activation function is selected. This activation function compares the values of multiple parallel convolution outputs and selects the maximum value at each position. This selection mechanism helps to improve the model's non-linear representation ability, thereby enhancing the model's performance in complex tasks.
[0186] In addition, in order to further enhance the network's expression ability and training effect, the present invention improves on the basis of the residual skip connection and adopts a gating mechanism to control the information flow.
[0187] The output of the skip connection is adjusted through a weighted operation, that is, the output is
[0188] g(x)×x+(1-g(x))×H(x)
[0189] where H(x) is the feature map processed through convolution operations, and g(x) is the weighting factor of the gating mechanism. This mechanism can dynamically adjust the fusion method of the input feature map and the feature map after convolution processing, thereby optimizing the information flow and feature selection.
[0190] The gating mechanism effectively prevents the overtransmission or loss of information by dynamically adjusting the activation degree of the residual connection, ensures the smooth transmission of information between different levels of the network, and at the same time avoids the problem of gradient disappearance in deep networks. In this way, the network can better retain important features during the training process and effectively accelerate the model convergence. The gating mechanism can add adaptive feature selection in the skip connection to help the network learn more meaningful feature representations at different levels.
[0191] Other steps and parameters are the same as those in the first or second specific implementation manner.
[0192] Specific implementation manner four: The difference between this implementation manner and one of the first to third specific implementation manners is that the deep learning network model is a pre-trained deep learning network model;
[0193] The pre-training process of the deep learning network model is as follows:
[0194] Input the image set with labeled categories and position information into the backbone network of the deep learning network model. The backbone network of the deep learning network model outputs feature vectors, and input the feature vectors of each image into the fully connected layer for classification;
[0195] Until the entire image set with labeled categories and position information is input into the backbone network, a pre-trained deep learning network model is obtained.
[0196] Other steps and parameters are the same as those in one of the first to third specific implementation manners.
[0197] Specific implementation manner five: The difference between this implementation manner and one of the first to fourth specific implementation manners is that in step two, when collecting video data, the rotation angle data of the camera is synchronously recorded; the specific process is as follows:
[0198] Use a gyroscope sensor to record the rotation angle data of the camera when shooting the video.
[0199] The gyroscope can measure the angular velocity of the camera in real time and calculate the rotation angle of the camera through integration.
[0200] Angle data synchronization: During the video acquisition process, ensure that the time stamps of the gyroscope sensor and the camera are synchronized, so as to accurately associate each frame of video with the corresponding rotation angle of the camera.
[0201] Angle data storage: Store the angular velocity data recorded by the gyroscope for subsequent calculation of the rotation angle of the camera when shooting key frames. These angle data will be used in the geometric constraint model to locate the position of the camera.
[0202] Other steps and parameters are the same as those in one of the first to fourth specific implementation manners.
[0203] Specific implementation manner six: The difference between this implementation manner and one of the first to fifth specific implementation manners is that in step three, key frame images are extracted from the video data based on the gray information entropy and combined with the time interval limit; the specific process is as follows:
[0204] Calculate the gray information entropy of each frame in the video data, sort all frames in descending order of entropy value, and obtain the sorted frame numbers;
[0205] Set the minimum frame interval threshold (assuming 60 frames per second, set the minimum frame interval threshold to 7 frames);
[0206] Delete the frame numbers with frame intervals less than the minimum frame interval threshold among the adjacent frame numbers in the sorted frame numbers, retain the frame numbers greater than or equal to the minimum frame interval threshold, and obtain the key frame images in the video data.
[0207] Other steps and parameters are the same as those in any one of the specific embodiments one to five.
[0208] Specific embodiment six or seven: The difference between this embodiment and any one of the specific embodiments one to six is that for the step of deleting the frame numbers with frame intervals less than the minimum frame interval threshold among the adjacent frame numbers in the sorted frame numbers and retaining the frame numbers greater than or equal to the minimum frame interval threshold to obtain the key frame images in the video data; the specific process is as follows:
[0209] The frame numbers of all frames sorted from high to low according to the entropy value are A, B, C, D, E, F, G;
[0210] If the frame interval between frame number A and frame number B is less than the minimum frame interval, discard frame number B, and continue to compare frame number A with frame number C until all frame numbers are traversed;
[0211] If the frame interval between frame number A and frame number B is greater than or equal to the minimum frame interval, retain frame number A and frame number B, and continue to compare frame number B with frame number C until all frame numbers are traversed;
[0212] Obtain the key frame images in the video data to ensure that the distribution of the key frames on the time axis is uniform and the information is rich.
[0213] After sorting, the original frame ordinal number of each frame remains unchanged (i.e., the serial number in the video). To avoid the situation where the selected key frames are too close and the information overlaps, a minimum frame interval is introduced. This process can be like this: The minimum frame interval threshold is 7 frames. The frames sorted from high to low according to the information entropy are 17, 14, 26 respectively; First, select 17. The interval between 14 and 17 is less than the minimum frame interval, so it is discarded. The interval between 26 and 17 is greater than the minimum frame interval, so it is selected;
[0214] Other steps and parameters are the same as those in any one of the specific embodiments one to six.
[0215] Specific embodiment eight: The difference between this embodiment and any one of the specific embodiments one to seven is that in step four, the similarity is the cosine value of the included angle between the feature vectors.
[0216] After the feature vectors are extracted, the feature vectors of the key frames are compared with those of each image in the image library to evaluate the semantic similarity between them. Semantic similarity refers to the degree of similarity in high-level semantic information of images, rather than simple pixel-level similarity. To quantify this semantic similarity, this study uses cosine similarity as the metric function. Cosine similarity measures the similarity between two feature vectors by calculating the cosine value of the angle between them, and can effectively capture the similarity in the direction of the feature vectors.
[0217] After calculating the similarity between the key frame and all images in the image library, three frames with the highest similarity are selected as the matching results on the premise of meeting the minimum frame interval constraint. These three frames represent the images in the image library that are semantically closest to the key frame and can effectively reflect the semantic content of the key frame. The matching results include not only the similarity values, but also the position information of the positioning points of the corresponding images in the image library. By inputting the coordinate and angle information of these three frames into the positioning algorithm and combining geometric constraints, the accurate positioning of the video shooting location is finally achieved.
[0218] Other steps and parameters are the same as those in any one of the first to seventh specific embodiments.
[0219] Specific embodiment nine: The difference between this embodiment and any one of the first to eighth specific embodiments is that in step five, based on the matching results of the key frames, the camera position is located on the map; the specific process is as follows:
[0220] The positions of the three key frames in the matching results (such as a store) are respectively marked as points A(x1, y1), B(x2, y2), and C(x3, y3) (points A(x1, y1), B(x2, y2), and C(x3, y3) in the plane rectangular coordinate system);
[0221] Through the angular velocity information recorded by the gyroscope sensor, the rotation angles of the camera when shooting adjacent key frames are calculated by integration, denoted as θ1 and θ2; the included angle is expressed through the relationship between vectors and points.
[0222] For example,
[0223] The included angle θ1 between the target point P(x, y) and the reference points A(x1, y1) and B(x2, y2) is calculated by the vector dot product formula:
[0224]
[0225] The included angle θ2 between the target point P(x, y) and the reference points B(x2, y2) and C(x3, y3) is calculated by the vector dot product formula:
[0226]
[0227] Among them,
[0228] The vector PA represents the line connecting the target point P(x, y) to the reference point A(x1, y1);
[0229] The vector PB represents the line connecting the target point P(x, y) to the reference point B(x2, y2);
[0230] The vector PC represents the line connecting the target point P(x, y) to the reference point C(x3, y3);
[0231] The target point P(x, y) is the camera position;
[0232] PA·PB represents the dot product of the vector PA and the vector PB;
[0233] PC·PB represents the dot product of the vector PC and the vector PB;
[0234] ||PA|| represents the magnitude of the vector PA;
[0235] ||PB|| represents the magnitude of the vector PB;
[0236] ||PC|| represents the magnitude of the vector PC;
[0237] According to the angle constraint, a non - linear equation system is obtained:
[0238]
[0239] Among them,
[0240] f1(x, y) represents the constraint condition of the angle between the vector PA and PB;
[0241] f2(x, y) represents the constraint condition of the angle between the vector PB and PC.
[0242] The non - linear equation system reflects the geometric constraints satisfied by the target point P(x, y). Solving P(x, y) can determine the position of the target point;
[0243] The Gauss - Newton method is used to iteratively solve the non - linear equation system to solve the position of the target point P(x, y); the specific process is as follows:
[0244] By initial value selection, Taylor expansion, least - squares method correction, and iterative update, the coordinates of the target point P are gradually optimized until the predetermined accuracy requirement is met.
[0245] Other steps and parameters are the same as those in any one of the first to eighth specific embodiments.
[0246] Specific embodiment ten: The difference between this embodiment and any one of the first to ninth specific embodiments is that the dot product of the vector PA and the vector PB is expressed as: PA·PB = (x1 - x)(x2 - x)+(y1 - y)(y2 - y);
[0247] The dot product of vector PC and vector PB is expressed as: PC·PB = (x3 - x)(x2 - x) + (y3 - y)(y2 - y);
[0248] The magnitude of vector PA, ||PA||, is expressed as:
[0249] The magnitude of vector PB, ||PB||, is expressed as:
[0250] The magnitude of vector PC, ||PC||, is expressed as:
[0251] Other steps and parameters are the same as those in any one of the first to ninth specific embodiments.
[0252] Due to possible slight errors in angle integration calculation and key frame matching, the present invention controls the angle error within 5 degrees through sensitivity analysis. Through multiple simulation experiments, the error range of target point positioning is verified, and finally a positioning result with centimeter-level accuracy is achieved.
[0253] In this step, through the geometric constraint model and the Gauss-Newton method, combined with the key frame matching result and the gyroscope angle information, the accurate positioning of the video shooting point is achieved, and the error is controlled within the decimeter level, meeting the requirements of high-precision indoor positioning.
[0254] The present invention may also have many other embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and deformations according to the present invention, but these corresponding changes and deformations should all fall within the protection scope of the appended claims of the present invention.
Claims
1. A video positioning method based on information entropy screening and feature point matching, characterized by: The specific process of the method is: Step 1: Construct a sample set of indoor images with unlabeled categories and labeled location information; Inputting an indoor image sample set with unlabeled categories and labeled location information into the backbone network of the deep learning network model, the backbone network of the deep learning network model outputs a feature vector, and obtaining a feature vector of each image in the indoor image sample set; Step 2: Video data acquisition, synchronously recording the camera's rotation angle data; Step 3: Extract key frame images from video data based on grayscale information entropy and combined with time interval restriction; Step 4: The key frame image extracted in step 3 is input into the backbone network of the deep learning network model, and the backbone network of the deep learning network model outputs the feature vector of the key frame image; Calculate the similarity between the feature vector of the i-th key frame image and the feature vector of each image in the indoor image sample set obtained in step 1; Select three key frames with the highest similarity to the feature vector of the i-th key frame image and the position information corresponding to the three key frames from the feature vector of each image in the indoor image sample set obtained in step 1 as the matching result of the i-th key frame image; Until the matching results of all key frame images are obtained; Step 5: Based on the matching results of the key frame images, locate the camera position in the map.
2. The video positioning method based on information entropy screening and feature point matching according to claim 1 is characterized in that: The deep learning network model includes: a backbone network and a fully connected layer; The backbone network includes: the first convolution layer, the second convolution layer, the third convolution layer, the fourth convolution layer, the fifth convolution layer, the sixth convolution layer, the Maxout activation function layer, the pooling layer, the seventh convolution layer, the eighth convolution layer, the ninth convolution layer, the tenth convolution layer, the eleventh convolution layer, the twelfth convolution layer, the thirteenth convolution layer, the fourteenth convolution layer, the fifteenth convolution layer, the sixteenth convolution layer, the seventeenth convolution layer, the eighteenth convolution layer, the nineteenth convolution layer, the twentieth convolution layer, the twenty-first convolution layer, the twenty-second convolution layer, the first gating mechanism, the second gating mechanism, the third gating mechanism, the fourth gating mechanism, the fifth gating mechanism, the sixth gating mechanism, the seventh gating mechanism, and the eighth gating mechanism; Each of the first gating mechanism, the second gating mechanism, the third gating mechanism, the fourth gating mechanism, the fifth gating mechanism, the sixth gating mechanism, the seventh gating mechanism, and the eighth gating mechanism comprises a 1×1 convolution layer and a Sigmoid activation function layer in sequence; The convolution kernel size of the first convolutional layer is 3×1 and the number of channels is 64; The convolution kernel size of the second convolutional layer is 5×1 and the number of channels is 64; The convolution kernel size of the third convolutional layer is 7×1 and the number of channels is 64; The convolution kernel size of the fourth convolutional layer is 1×3 and the number of channels is 64; The convolution kernel size of the fifth convolutional layer is 1×5 and the number of channels is 64; The convolution kernel size of the sixth convolutional layer is 1×7 and the number of channels is 64; The convolution kernel size of the seventh convolutional layer is 3×3, the number of channels is 64, and the stride is 2; The convolution kernel size of the eighth convolutional layer is 3×3, the number of channels is 64, and the stride is 2; The convolution kernel size of the ninth convolutional layer is 3×3, the number of channels is 64, and the stride is 2; The convolution kernel size of the tenth convolutional layer is 3×3, the number of channels is 64, and the stride is 2; The convolution kernel size of the eleventh convolutional layer is 3×3, the number of channels is 128, and the stride is 2; The convolution kernel size of the twelfth convolutional layer is 3×3, the number of channels is 128, and the stride is 2; The convolution kernel size of the thirteenth convolutional layer is 3×3, the number of channels is 128, and the stride is 2; The convolution kernel size of the fourteenth convolutional layer is 3×3, the number of channels is 128, and the stride is 2; The convolution kernel size of the fifteenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2; The convolution kernel size of the sixteenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2; The convolution kernel size of the seventeenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2; The convolution kernel size of the eighteenth convolutional layer is 3×3, the number of channels is 256, and the stride is 2; The convolution kernel size of the nineteenth convolutional layer is 3×3, the number of channels is 512, and the stride is 2; The convolution kernel size of the twentieth convolutional layer is 3×3, the number of channels is 512, and the stride is 2; The convolution kernel size of the 21st convolutional layer is 3×3, the number of channels is 512, and the stride is 2; The convolution kernel size of the 22nd convolutional layer is 3×3, the number of channels is 512, and the stride is 2.
3. The video positioning method based on information entropy screening and feature point matching according to claim 2 is characterized in that: The working process of the deep learning network model is: The image is input into the first convolutional layer, the second convolutional layer, and the third convolutional layer of the deep learning network model respectively. The first convolutional layer outputs feature α, the second convolutional layer outputs feature α′, and the third convolutional layer outputs feature α″; The output feature α of the first convolutional layer is input into the fourth convolutional layer, and the fourth convolutional layer outputs the feature α″′; The second convolutional layer outputs the feature α′ which is input into the fifth convolutional layer. The fifth convolutional layer outputs the feature The output feature α″ of the third convolutional layer is input into the sixth convolutional layer, and the output feature The fourth convolutional layer outputs the feature α″′, and the fifth convolutional layer outputs the feature The sixth convolutional layer outputs features Input Maxout activation function layer, Maxout activation function layer output features α ; Maxout activation function layer output features α Input the pooling layer, and the pooling layer outputs feature β; The output feature β of the pooling layer is input into the seventh convolutional layer and the eighth convolutional layer in sequence, and the eighth convolutional layer outputs the feature β′; The output feature β of the pooling layer is input into the first gating mechanism, and the first gating mechanism outputs the feature β″; The output feature β′ of the eighth convolutional layer and the output feature β″ of the first gating mechanism are input into the ninth convolutional layer and the tenth convolutional layer in sequence, and the output feature β″′ of the tenth convolutional layer is output; The output feature β′ of the eighth convolutional layer and the output feature β″ of the first gating mechanism are input into the second gating mechanism, and the output feature of the second gating mechanism is The output feature β″′ of the tenth convolutional layer and the output feature of the second gating mechanism are Input the eleventh convolution layer and the twelfth convolution layer in sequence, and the twelfth convolution layer outputs the features The tenth convolutional layer outputs the feature β″′ and the second gating mechanism outputs the feature Input the third gating mechanism, the third gating mechanism outputs features β ; The output features of the twelfth convolutional layer And the third gating mechanism output characteristics β Input the 13th convolutional layer and the 14th convolutional layer in sequence, and the 14th convolutional layer outputs the feature γ; The output features of the twelfth convolutional layer And the third gating mechanism output characteristics β Input the fourth gating mechanism, and the fourth gating mechanism outputs the feature γ′; The output feature γ of the fourteenth convolutional layer and the output feature γ′ of the fourth gating mechanism are input into the fifteenth convolutional layer and the sixteenth convolutional layer in sequence, and the output feature γ″ of the sixteenth convolutional layer is obtained; The output feature γ of the fourteenth convolutional layer and the output feature γ′ of the fourth gating mechanism are input into the fifth gating mechanism, and the fifth gating mechanism outputs the feature γ″′; The output feature γ″ of the sixteenth convolutional layer and the output feature γ″′ of the fifth gating mechanism are input into the seventeenth convolutional layer and the eighteenth convolutional layer in sequence. The output feature of the eighteenth convolutional layer is The output feature γ″ of the sixteenth convolutional layer and the output feature γ″′ of the fifth gating mechanism are input into the sixth gating mechanism, and the output feature The output features of the eighteenth convolutional layer And the sixth gating mechanism output characteristics Input the 19th convolutional layer and the 20th convolutional layer in sequence, and the 20th convolutional layer outputs the features γ ; The output features of the eighteenth convolutional layer And the sixth gating mechanism output characteristics Input the seventh gating mechanism, and the seventh gating mechanism outputs feature δ; The twentieth convolutional layer output features γ The output feature δ of the seventh gating mechanism is sequentially input into the twenty-first convolutional layer and the twenty-second convolutional layer, and the twenty-second convolutional layer outputs the feature δ′; The output features of the twentieth convolutional layer γ The seventh gating mechanism outputs feature δ which is input into the eighth gating mechanism, and the eighth gating mechanism outputs feature δ″; The feature vector obtained by concatenating the output feature δ′ of the 22nd convolutional layer and the output feature δ″ of the 8th gating mechanism is the output feature vector of the backbone network in the deep learning network model; The backbone network outputs the feature vector which is input into the fully connected layer, and the fully connected layer outputs the classification result.
4. The video positioning method based on information entropy screening and feature point matching according to claim 3 is characterized in that: The deep learning network model is a pre-trained deep learning network model; The pre-training process of the deep learning network model is: The image set with labeled category and location information is input into the backbone network in the deep learning network model. The backbone network in the deep learning network model outputs a feature vector. The feature vector of each image is input into the fully connected layer for classification. Until all the image sets with labeled categories and location information are input into the backbone network, a pre-trained deep learning network model is obtained.
5. The video positioning method based on information entropy screening and feature point matching according to claim 4 is characterized in that: The video data acquisition in step 2 synchronously records the rotation angle data of the camera; the specific process is: Use the gyroscope sensor to record the rotation angle data of the camera when shooting video.
6. The video positioning method based on information entropy screening and feature point matching according to claim 5 is characterized in that: In the step 3, based on the grayscale information entropy and combined with the time interval restriction, the key frame image in the video data is extracted; The specific process is: Calculate the grayscale information entropy of each frame in the video data, sort all frames from high to low according to the entropy value, and obtain the sorted frame number; Set the minimum frame interval threshold; Delete the frame numbers whose frame intervals between adjacent frame numbers in the sorted frame numbers are less than the minimum frame interval threshold, retain the frame numbers that are greater than or equal to the minimum frame interval threshold, and obtain the key frame images in the video data.
7. The video positioning method based on information entropy screening and feature point matching according to claim 6 is characterized in that: The frame numbers whose frame intervals between adjacent frame numbers in the sorted frame numbers are less than the minimum frame interval threshold are deleted, and the frame numbers whose frame intervals are greater than or equal to the minimum frame interval threshold are retained to obtain the key frame images in the video data; the specific process is: The frame numbers after all frames are sorted from high to low according to entropy value are A, B, C, D, E, F, G; If the frame interval between frame number A and frame number B is less than the minimum frame interval, frame number B is discarded, and frame number A is compared with frame number C until all frame numbers are traversed; If the frame interval between frame number A and frame number B is greater than or equal to the minimum frame interval, retain frame number A and frame number B, and continue to compare frame number B with frame number C until all frame numbers are traversed; Get the key frame image in the video data.
8. The video positioning method based on information entropy screening and feature point matching according to claim 7 is characterized in that: The similarity in step 4 is to calculate the cosine value of the angle between the feature vectors.
9. The video positioning method based on information entropy screening and feature point matching according to claim 8 is characterized in that: In step 5, the camera position is located in the map based on the key frame matching result; the specific process is: Mark the positions of the three key frames in the matching results as points A(x1,y1), B(x2,y2), and C(x3,y3); The angle θ1 between the target point P(x,y) and the reference points A(x1,y1) and B(x2,y2) is calculated by the vector dot product formula: The angle θ2 between the target point P(x,y) and the reference points B(x2,y2) and C(x3,y3) is calculated by the vector dot product formula: in, Vector PA represents the line from the target point P(x,y) to the reference point A(x1,y1); Vector PB represents the line from the target point P(x,y) to the reference point B(x2,y2); Vector PC represents the line from the target point P(x,y) to the reference point C(x3,y3); The target point P(x,y) is the camera position; PA·PB represents the dot product of vector PA and vector PB; PC·PB represents the dot product of vector PC and vector PB; ||PA|| represents the modulus of the vector PA; ||PB|| represents the modulus of vector PB; ||PC|| represents the modulus length of vector PC; According to the angle constraint, we get the nonlinear equations: in, f1(x,y) represents the constraint condition of the angle between vectors PA and PB; f2(x,y) represents the constraint condition on the angle between vectors PB and PC. The Gauss-Newton method is used to iteratively solve the nonlinear equations to solve the target point position P(x,y).
10. The video positioning method based on information entropy screening and feature point matching according to claim 9 is characterized in that: The dot product of the vector PA and the vector PB is expressed as: PA·PB=(x1-x)(x2-x)+(y1-y)(y2-y); The dot product of vector PC and vector PB is expressed as: PC·PB=(x3-x)(x2-x)+(y3-y)(y2-y); The modulus length ||PA|| of the vector PA is expressed as: The modulus length ||PB|| of the vector PB is expressed as: The modulus length ||PC|| of the vector PC is expressed as: