Target Location Method, System and Storage Medium Based on Multi-Source Data Fusion
Through the fusion of camera images and sound information, the problem of insufficient accuracy in traditional indoor positioning technology is solved, and more accurate target positioning is achieved.
Patent Information
- Application Number
- CN202111498779.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-09
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-12-09
AI Technical Summary
The existing indoor positioning technology relies on single data for positioning, which has the problem of accuracy, especially in complex environments, and it is difficult to meet high-precision requirements. Traditional camera positioning is susceptible to occlusion and light changes.
Multi-source data fusion is carried out in combination with camera image information and sound information. After initial positioning through image recognition, sound information is corrected and optimized, and precise positioning is finally obtained.
The positioning accuracy is improved and the shortcomings of single data positioning are overcome, especially in the case of target image occlusion or scale uncertainty, and higher positioning accuracy is achieved.
Smart Images

Figure CN114119755B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of target recognition and tracking, and particularly to a target positioning method, system and storage medium based on multi-source data fusion. Background Art
[0002] In recent years, indoor location services with increasingly urgent demands have become a research hotspot in the mobile Internet era. The indoor positioning technology has developed rapidly and gradually plays a role in all walks of life, bringing certain beneficial impacts to people's daily lives. With the application and development of related technologies based on user location information, location services have become a basic service demand necessary for people's daily work and life. Especially in large and complex indoor environments, such as museums, airports, supermarkets, hospitals, large offices and other areas, people have an urgent need for location services. Driven by the rapid development of the mobile Internet and the application demand of location services, the current indoor positioning technology is in a relatively fast development stage, and researchers have proposed many theories and methods of indoor positioning technology.
[0003] Currently, image-based indoor positioning technology mainly uses a monocular camera or a binocular camera. It is difficult to establish the corresponding relationship with coordinates and sizes from a single picture taken by a monocular camera, and there is scale uncertainty. At the same time, a monocular camera cannot directly obtain the depth information of the objects in the image.
[0004] A binocular camera takes pictures at the same time through two cameras on the left and right. Since the two pictures are taken of the same object from different angles, the distance can be measured through the parallax of the two pictures. The closer a point is to the camera, the greater its parallax in the left and right cameras; the farther a point is from the image plane, the smaller its parallax in the left and right cameras. Therefore, using a binocular camera for target positioning has a better effect than a monocular camera in personnel positioning and monitoring.
[0005] However, traditional camera image positioning methods are easily affected by factors such as the small coverage area of the camera, the signal being easily blocked by buildings or intentionally blocked by monitored personnel, light changes, large data transmission and processing volume, etc. When the camera captures different sides of the same person, it is also difficult to accurately identify, and single distance measurement still cannot meet the high-precision positioning requirements. It is difficult to achieve the positioning requirements of the target person only relying on image positioning. Existing technologies have accuracy problems when relying on single data for positioning, so the application of indoor positioning technology needs to be further expanded. Summary of the Invention
[0006] The present invention aims to provide a target positioning method based on multi-source data fusion. On the basis of identifying target positioning by collecting images with a camera, the sound information in the target direction collected is used to further correct and optimize the target positioning, so as to obtain more accurate positioning, and the positioning effect is better than that of traditional image recognition.
[0007] The technical solution provided by the present invention is: a target positioning method based on multi-source data fusion, including:
[0008] S1: Enter the scene information of the monitoring area;
[0009] S2: Identify whether a target enters the monitoring area;
[0010] S3: Collect the image information of the target in the monitoring area, combine it with the scene information of the monitoring area, judge the position where the target is located, and obtain the first positioning of the target;
[0011] S4: Collect the sound information in the direction of the target in the monitoring area, identify the sound content, combine it with the scene information of the monitoring area, judge the position where the target is located, and obtain the second positioning of the target;
[0012] S5: Correct the first positioning according to the information of the second positioning to obtain an accurate positioning.
[0013] The working principle and advantages of the present invention are as follows: First, the scene information of the monitoring area is entered into the system to facilitate subsequent analysis and recognition. After identifying that a target enters the monitoring area, the system collects the image information of the target in the monitoring area, combines it with the scene information of the monitoring area, and performs the first positioning on the position where the target is located. Due to reasons such as the target image being too far away, the target image being blocked, or scale uncertainty, this positioning is only a relatively broad positioning through image recognition. At the same time, the sound information in the direction of the target in the monitoring area is collected, the sound content is identified, and according to the sound content combined with the scene information of the monitoring area, the second positioning of the target is performed. After obtaining the two positioning information, on the basis of the first positioning, the information of the second positioning is used to correct and optimize the first positioning, further narrowing the range of the first positioning to obtain a relatively accurate positioning. Compared with pure camera image positioning, the present invention introduces the dimension of sound information to correct and optimize the positioning accuracy, overcomes the disadvantage of low positioning accuracy caused by reasons such as the target image being too far away, the target image being blocked, or scale uncertainty in pure image positioning, and achieves a better positioning effect.
[0014] Further, the scene information of the monitoring area entered in S1 includes the vector map and building structure of the monitoring area.
[0015] The target positioning method of the present invention is mainly used for indoor target positioning. The scene information of the monitoring area entered includes vector map and building structure information. The vector map is a 2D or 3D map, which is used to understand the walking route in the monitoring area. The building structure includes information such as doors, support columns, or walls in the monitoring area, which is used to understand the entrances, exits, and occlusion structures in the monitoring area.
[0016] Further, the identification method in S2 is thermal imaging identification.
[0017] Thermal imaging recognition can quickly detect the entry of a target into the monitoring area and is suitable as the first step of recognition for target positioning.
[0018] Furthermore, the image information of the target in the monitoring area collected in S3 includes image information in two or more directions.
[0019] Before performing sound recognition, collecting target images for positioning using two or more cameras can obtain a relatively accurate target image positioning, facilitating subsequent further positioning. The basic principle is that a triangle is formed among the two cameras and the target. The distance between the two cameras and the rotation angle of the cameras are known, and the perpendicular distance from the target to the line where the two cameras are located can be calculated using the triangle principle with formulas.
[0020] Furthermore, S3 includes:
[0021] S3-1: Real-time collect the image information of the target in the monitoring area;
[0022] S3-2: Based on the real-time collected image information, obtain the historical movement trajectory of the target;
[0023] S3-3: Based on the historical movement trajectory of the target and combined with the scene information of the monitoring area, obtain the predicted movement trajectory of the target;
[0024] S3-4: Based on the real-time collected image information, find the position closest to the real-time position on the trajectory as the first positioning.
[0025] By analyzing the images collected before and after the target, first obtain the historical movement trajectory of the target, match the historical movement trajectory with the scene information of the monitoring area, such as the walking route map in the room, to obtain the route with the highest similarity, and based on this route, the predicted movement trajectory of the target in the next step can be inferred. Based on the predicted movement trajectory and combined with the real-time collected target image, directly perform positioning on the predicted movement trajectory, which reduces unnecessary operation and analysis processes and can obtain more accurate positioning information.
[0026] Furthermore, S3 includes:
[0027] S3-1: Real-time collect the image information of the target in the monitoring area;
[0028] S3-2: Based on the real-time collected image information, through face recognition, obtain the identity of the target person;
[0029] S3-3: Obtain the historical movement trajectory corresponding to the target person and set it as the predicted movement trajectory;
[0030] S3-4: Based on the real-time collected image information, use the real-time position of the target person on the predicted movement trajectory as the first positioning.
[0031] The system records the locations where a person who often moves within the monitoring area frequently stays and the trajectories they often pass through. When the target entering the monitoring area is identified as a recorded target person, the system can roughly infer the possible target locations they may go to and set the route to the target locations as the predicted movement trajectory. Based on the predicted movement trajectory and combined with the real-time collected target images, positioning is directly carried out on the predicted movement trajectory, reducing unnecessary operation and analysis processes and obtaining relatively accurate positioning information.
[0032] Further, S4 includes:
[0033] S4-1: Collect the sound information in the direction of the target within the monitoring area and identify the sound content;
[0034] S4-2: If it is identified that the sound content includes the noises during the target's movement, then judge the direction of the source of the noises;
[0035] S4-3: According to the direction of the source of the noises and combined with the scene information of the monitoring area, judge the position where the target is located to obtain the second positioning of the target.
[0036] If the identified sound content is only the noises during the target's movement, such as the sound of opening or closing a door, footsteps, and the sound of collision with an object, the system then judges the position where the target is located according to the direction of the source of the above-mentioned sound noises and combined with the scene information of the monitoring area, and the position information of the target can be obtained only through the sound information to obtain the second positioning.
[0037] Further, S4 includes:
[0038] S4-1: Collect the sound information in the direction of the target within the monitoring area and identify the sound content;
[0039] S4-2-1: If it is identified that the sound content includes the noises during the target's movement, then judge the direction of the source of the noises;
[0040] S4-2-2: The identified sound content also includes keywords related to the scene information of the monitoring area. According to the keywords and combined with the scene information of the monitoring area, judge the target locations near the target;
[0041] S4-3-1: Set the route from the target to the target locations as the predicted movement trajectory;
[0042] S4-3-2: According to the direction of the source of the noises and combined with the scene information of the monitoring area, judge the position where the target is located and use the real-time position of the target on the predicted movement trajectory as the second positioning.
[0043] If the recognized sound also includes the target's speech content, and the speech content includes keywords related to the scene information of the monitoring area, such as a conference room, a turn, etc., the system will analyze the scene around the target that corresponds to the keyword based on the keyword, set it as the target location, and set the route from the target to the target location as the predicted movement trajectory based on the map of the monitoring area. Combined with the direction of the sound source during the target's movement, real-time positioning is performed on the predicted movement trajectory.
[0044] The present invention also provides a target positioning system based on multi-source data fusion, which adopts the target positioning method based on multi-source data fusion.
[0045] The present invention also provides a target positioning storage medium based on multi-source data fusion, which is used to store computer executable instructions. When the computer executable instructions are executed, the target positioning method based on multi-source data fusion is realized. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 This is a logic block diagram of Embodiment 1 of the target positioning method based on multi-source data fusion of the present invention. DETAILED DESCRIPTION
[0047] Embodiment 1:
[0048] This embodiment discloses a target positioning method based on multi-source data fusion, and the logical flow is as follows: Figure 1 As shown, the method embodiment includes the following steps (the numbering of each step in this solution is only used to distinguish the steps, and does not limit the specific execution order of each step. In some cases, each step can be performed simultaneously):
[0049] S1: inputting scene information of the monitoring area. In this embodiment, what is inputted is the completed design drawing of the monitoring area, including the floor plan and building structure drawing of the monitoring area.
[0050] S2: Identify whether there is a target entering the monitoring area, and use the thermal imaging recognition device to identify whether there is anyone entering the monitoring area.
[0051] S3: Collect the image information of the target in the monitoring area, combine it with the scene information of the monitoring area, judge the position where the target is located, and obtain the first positioning of the target. There are two cameras set in the monitoring area, and the cameras are directly aimed at the monitoring target. The two cameras respectively collect the image information of the target. According to the deflection angles of the two cameras relative to the line connecting the cameras, as well as the distance between the two cameras and the deflection angles, the actual distances of the target from each camera are obtained. According to the actual distances of the target from each camera and the deflection angles of the line connecting the target and the camera relative to the horizontal plane where the camera is located, the vertical distance of the target from the horizontal plane where the camera is located is obtained, so as to roughly obtain the position information where the target is located, and obtain its position as the first positioning. The target is walking on the office aisle, but specifically which workbench it reaches is not recognized and located by the system because the target walks too far away from the camera and enters the visual blind area, and part of the image is blocked.
[0052] S4-1: Collect the sound information of the target direction in the monitoring area and identify the sound content.
[0053] S4-2: If it is recognized that the sound content includes the noise during the movement of the target, then judge the direction of the noise source. In this embodiment, the recognized sound information is only the door opening sound and the walking sound of the target entering the monitoring area. The noise sound is collected through the microphone array set in the monitoring area. According to the principle of sound source localization of the microphone array, the multi-channel sound signals collected are processed to obtain the sound source localization of the noise.
[0054] S4-3: According to the direction of the noise source, combine it with the scene information of the monitoring area, judge the position where the target is located, and obtain the second positioning of the target. After obtaining the localization of the noise sound source, combined with the floor plan layout in the monitoring area, it is thus judged that the target's footsteps are getting closer and closer to the No. 7 workbench and finally stop at the No. 7 workbench. Therefore, this position is obtained as the second positioning, and the target is located near the No. 3 workbench.
[0055] S5: Correct the first positioning according to the information of the second positioning to obtain the precise positioning. The final position of the first positioning is at the turning point of the horizontal row where the No. 7 workbench is located, and according to the footsteps of the target stopping near the No. 7 workbench in the second positioning, it is obtained that the position where the target is located is the positioning point of the No. 7 workbench.
[0056] Embodiment 2:
[0057] This method embodiment includes the following steps (the numbers of each step in this solution are only used for step distinction, do not limit the specific execution order of each step, and each step can also be carried out simultaneously):
[0058] S1: Enter the scene information of the monitoring area. In this embodiment, the entered is the completion design drawing of the monitoring area, including the floor plan layout and building structure drawing of the monitoring area.
[0059] S2: Identify whether a target enters the monitoring area. Use a thermal imager to identify whether a person enters the monitoring area.
[0060] S3-1: Collect the image information of the target in the monitoring area in real time, that is, the real-time image information starting from when the target enters the monitoring area.
[0061] S3-2: Based on the image information collected in real time, obtain the historical movement trajectory of the target. According to the image information from when the target enters the monitoring area and its subsequent activities, concatenate the image position information of the target at different time points to obtain the historical movement trajectory of the target.
[0062] S3-3: Based on the historical movement trajectory of the target and combined with the scene information of the monitoring area, obtain the predicted movement trajectory of the target. That is, according to the route the target has taken and combined with the possible routes that can be continued on the floor plan of the monitoring area, it is used as the predicted movement trajectory. In this embodiment, there are two predicted movement trajectories here.
[0063] S3-4: Based on the image information collected in real time, find the position closest to the real-time position on the trajectory as the first positioning. Based on the predicted movement trajectory, locate the real-time position of the target according to the image positioning of the dual-camera system in the monitoring area as the first positioning.
[0064] S4-1: Collect the sound information of the target direction in the monitoring area and identify the sound content.
[0065] S4-2: If it is recognized that the sound content includes the noise during the movement of the target, then judge the direction of the source of the noise. When the footsteps of the target are recognized, analyze the direction of the source of the noise.
[0066] S4-3: Based on the direction of the source of the noise and combined with the scene information of the monitoring area, judge the position where the target is located to obtain the second positioning of the target. Based on the direction of the source of the noise and combined with the floor plan of the monitoring area, judge the position where the target has walked to as the second positioning.
[0067] S5: Correct the first positioning according to the information of the second positioning to obtain the accurate positioning. According to the sound positioning, judge which predicted movement trajectory the noise of the target is closer to, and through image positioning and sound positioning, display the accurate positioning of the target on this predicted movement trajectory in real time.
[0068] Embodiment Three:
[0069] This method embodiment includes the following steps (in this solution, the numbers of each step are only used for step distinction, do not limit the specific execution order of each step, and each step can also be carried out simultaneously):
[0070] S1: Input the scene information of the monitoring area. In this embodiment, the input is the as-built design drawing of the monitoring area, including the floor plan and building structure diagram of the monitoring area.
[0071] S2: Identify whether there is a target entering the monitoring area. Use a thermal imaging detector to identify whether anyone enters the monitoring area.
[0072] S3-1: Collect the image information of the target in the monitoring area in real time. When the target is closest to the camera, collect the face image of the target.
[0073] S3-2: Based on the image information collected in real time, identify the identity of the target person through face recognition. Compare the collected face image of the target with the database to identify that the target is an employee of the company.
[0074] S3-3: Obtain the historical movement trajectory corresponding to the target person and set it as the predicted movement trajectory. Extract the historical movement trajectory data of this employee from the database, and use the historical trajectory with the highest walking frequency of this employee as the predicted movement trajectory for this image positioning. That is, in this monitoring area, the historical trajectory with the highest walking frequency of this employee is the route from the workstation to the director's office.
[0075] S3-4: Based on the image information collected in real time, use the real-time position of the target person on the predicted movement trajectory as the first positioning. Based on the predicted movement trajectory, and according to the image positioning of the dual-camera system in the monitoring area, locate the real-time position of the employee as the first positioning.
[0076] S4-1: Collect the sound information of the target direction in the monitoring area and identify the sound content.
[0077] S4-2-1: If it is identified that the sound content includes the noise during the movement of the target, then judge the direction of the noise source. When the footsteps of the target are identified, analyze the direction of the noise source.
[0078] S4-2-2: The identified sound content also includes keywords related to the scene information of the monitoring area. According to the keywords and combined with the scene information of the monitoring area, judge the target location near the target. In this embodiment, it is identified that the employee mentions keywords such as "director" and "report" during the conversation. According to the scene layout in the monitoring area, judge that the director's office is the target location.
[0079] S4-3-1: Set the route from the target to the target location as the predicted movement trajectory. Set the shortest route from the current position of this employee to the director's office as the predicted movement trajectory.
[0080] S4-3-2: Based on the direction of the source of the sound, combined with the scene information of the monitoring area, judge the location of the target, and use the real-time position of the target on the predicted movement trajectory as the second positioning. Through the recognition of the direction of the footsteps of the microphone array on the employee's walking path, it is recognized that the sound of the employee's footsteps is getting closer and closer to the director's office, and the position where the employee is currently walking is judged as the second positioning.
[0081] S5: Correct the first positioning according to the information of the second positioning to obtain an accurate positioning. The paths of the first positioning and the second positioning basically coincide. Through image positioning and sound positioning, the accurate positioning of the target on the predicted movement trajectory is displayed in real time.
[0082] This embodiment also includes a target positioning system based on multi-source data fusion, and this system adopts the above-mentioned target positioning method based on multi-source data fusion.
[0083] This embodiment also includes a target positioning storage medium based on multi-source data fusion, which is used to store computer-executable instructions, and the computer-executable instructions, when executed, implement the above-mentioned target positioning method based on multi-source data fusion.
[0084] The above are only embodiments of the present invention. Specific structures and common knowledge such as characteristics that are well known in the art are not described in detail here. Those of ordinary skill in the art know all the common technical knowledge in the technical field to which the invention belongs before the application date or the priority date, can know all the existing technologies in this field, and have the ability to apply the conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration obtained from this application, combine their own abilities to improve and implement this solution. Some typical well-known structures or well-known methods should not become an obstacle for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can still be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent. The protection scope required by this application should be subject to the content of its claims, and the specific implementation manners and the like recorded in the specification can be used to interpret the content of the claims.
Claims
1. A target positioning method based on multi-source data fusion, characterized in that, Including: S1: Input the scene information of the monitoring area; S2: Identify whether there is a target entering the monitoring area; S3: Collect the image information of the target in the monitoring area, combine it with the scene information of the monitoring area, judge the location where the target is located, and obtain the first positioning of the target; S4: Collect the sound information in the direction of the target in the monitoring area, identify the sound content, combine it with the scene information of the monitoring area, judge the location where the target is located, and obtain the second positioning of the target; S5: Correct the first positioning according to the information of the second positioning to obtain the accurate positioning; The said S3 includes: S3-1: Collect the image information of the target in the monitoring area in real time; S3-2: According to the image information collected in real time, obtain the identity of the target person through face recognition; S3-3: Obtain the historical movement trajectory corresponding to the target person and set it as the predicted movement trajectory; S3-4: According to the image information collected in real time, take the real-time position of the target person on the predicted movement trajectory as the first positioning; The said S4 includes: S4-1: Collect the sound information in the direction of the target in the monitoring area and identify the sound content; S4-2-1: If it is recognized that the sound content includes the noise during the movement of the target, judge the direction of the source of the noise; S4-2-2: The recognized sound content also includes keywords related to the scene information of the monitoring area. According to the keywords and combined with the scene information of the monitoring area, judge the target location near the target; S4-3-1: Set the route from the target to the target location as the predicted movement trajectory; S4-3-2: According to the direction of the source of the noise, combined with the scene information of the monitoring area, judge the location where the target is located, and take the real-time position of the target on the predicted movement trajectory as the second positioning.
2. The target positioning method based on multi-source data fusion according to claim 1, characterized in that: The scene information input in the said S1 includes the vector map and building structure of the monitoring area.
3. The target positioning method based on multi-source data fusion according to claim 1, characterized in that: The recognition method in the said S2 is thermal imaging recognition.
4. The target positioning method based on multi-source data fusion according to claim 1, characterized in that: The image information of the target in the monitoring area collected in the said S3 includes image information in two or more directions.
5. The target positioning method based on multi-source data fusion according to claim 1, characterized in that: The said S3 includes: S3-1: Collect the image information of the target in the monitoring area in real time; S3-2: According to the image information collected in real time, obtain the historical movement trajectory of the target; S3-3: According to the historical movement trajectory of the target, combined with the scene information of the monitoring area, obtain the predicted movement trajectory of the target; S3-4: According to the image information collected in real time, find the position closest to the real-time position on the trajectory as the first positioning.
6. The target positioning method based on multi-source data fusion according to claim 1, wherein: The said S4 includes: S4-1: Collect the sound information in the direction of the target in the monitoring area and identify the sound content; S4-2: If it is recognized that the sound content includes the noise during the movement of the target, judge the direction of the source of the noise; S4-3: According to the direction of the source of the noise, combined with the scene information of the monitoring area, judge the location where the target is located, and obtain the second positioning of the target.
7. A target positioning system based on multi-source data fusion, characterized in that: This system adopts the target positioning method based on multi-source data fusion described in any one of claims 1 to 6.
8. A target positioning storage medium based on multi-source data fusion, which is used to store computer-executable instructions, and is characterized in that: The computer-executable instructions, when executed, implement the target positioning method based on multi-source data fusion described in any one of claims 1-6 above.
Citation Information
Patent Citations
Sound source locating method and device and target snapshot system
CN109683135A
A trajectory prediction method and related equipment
CN112805730A