Scene distinguishing method, device and equipment, sweeper and computer program product
By using the visual images of specific targets collected by the mobile device, determining their location information and matching them with the scene map, the problem of inability to effectively distinguish similar scenes in the prior art is solved, and the accurate positioning and functional stability of the mobile device are achieved.
Patent Information
- Application Number
- CN202510090637.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art cannot effectively distinguish similar scenarios, causing the mobile device positioning results to jump back and forth between similar scenarios, affecting the overall function.
By determining the location information of a specific target based on the visual image of a specific target collected by the mobile device, the location information of a specific target is determined, and using this information to match the scene map, the scene in which the specific target is located is determined, and finally determining the scene in which the mobile device is located is determined based on the scene in which the specific target is located.
A detailed distinction between similar scenes is achieved, and the positioning results are avoided from jumping back and forth between similar scenes, ensuring that the overall function of the mobile device is not affected.
Smart Images

Figure CN120093180A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of scene differentiation, and in particular to a scene differentiation method, device, equipment, sweeper and computer program product. Background Art
[0002] Mobile devices need to determine the correct posture according to their location in the scene to ensure that they can complete the predetermined task. However, current mobile devices cannot effectively distinguish similar scenes, causing the positioning results to jump back and forth between similar scenes, affecting the overall function of the mobile device. Therefore, how to distinguish similar scenes has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention
[0003] In view of this, the present application proposes a scene distinguishing method, device, equipment, sweeper and computer program product to solve the problem in the prior art that mobile devices cannot effectively distinguish similar scenes, causing the positioning results to jump back and forth between similar scenes, affecting the overall function of the mobile device.
[0004] The technical solutions proposed in this application are as follows:
[0005] In a first aspect, an embodiment of the present application provides a method for distinguishing scenes, including:
[0006] Determine the location information of the specific target based on the visual image of the specific target collected by the mobile device; the specific target includes a target with a fixed location;
[0007] Using the location information of the specific target, matching the specific target with a scene map is performed to determine the scene where the specific target is located; the scene map is generated based on the location information of the specific target in each scene;
[0008] The scene where the mobile device is located is determined according to the scene where the specific target is located.
[0009] Furthermore, in the above method, determining the location information of the specific target based on the visual image of the specific target collected by the mobile device includes:
[0010] Based on a visual image of a specific target captured by a mobile device, coordinates of a grounding line of the specific target are determined as position information of the specific target; the grounding line of the specific target includes a line formed when the specific target contacts the ground.
[0011] Furthermore, in the above method, the step of determining the coordinates of the grounding line of the specific target as the location information of the specific target based on the visual image of the specific target collected by the mobile device includes:
[0012] Determining image pixel coordinates of a ground line of the specific target based on a visual image of the specific target captured by the mobile device;
[0013] Perform an inverse perspective transformation on the image pixel coordinates of the grounding line of the specific target to obtain the coordinates of the grounding line of the specific target in a bird's-eye view, and determine the coordinates of the grounding line of the specific target in a bird's-eye view as the position information of the specific target.
[0014] Furthermore, in the above method, determining the scene where the mobile device is located according to the scene where the specific target is located includes:
[0015] Matching the point cloud map with the radar point cloud collected by the mobile device to obtain a scene matching result; the point cloud map is generated based on the radar point cloud in each scene;
[0016] The scene where the mobile device is located is determined according to the scene matching result and the scene where the specific target is located.
[0017] Furthermore, if it is determined that the scene where the mobile device is located includes at least two scenes according to the scene where the specific target is located, the above method further includes:
[0018] Determining three-dimensional information of the specific target based on a visual image of the specific target acquired by the mobile device;
[0019] Using the three-dimensional information of the specific target, matching the specific target with the three-dimensional map to determine the scene where the specific target is located; the three-dimensional map is generated based on the three-dimensional information of the specific target in each scene;
[0020] The scene where the mobile device is located is determined according to the scene where the specific target is located.
[0021] Furthermore, in the above method, determining the three-dimensional information of the specific target based on the visual image of the specific target collected by the mobile device includes:
[0022] Determining depth information of the specific target based on a visual image of the specific target captured by the mobile device;
[0023] The coordinates of the specific target in the three-dimensional space are determined as the three-dimensional information of the specific target by using the depth information of the specific target and the coordinates of the mobile device in the three-dimensional space.
[0024] In a second aspect, an embodiment of the present application provides a scene distinguishing device, including:
[0025] A first determining unit is used to determine the location information of a specific target based on a visual image of the specific target collected by a mobile device; the specific target includes a target with a fixed position;
[0026] A matching unit, used to match the specific target with a scene map by using the location information of the specific target, and determine the scene where the specific target is located; the scene map is generated according to the location information of the specific target in each scene;
[0027] The second determining unit is used to determine the scene where the mobile device is located according to the scene where the specific target is located.
[0028] In a third aspect, an embodiment of the present application provides an electronic device, including:
[0029] A memory and a processor; wherein the memory is used to store programs; and the processor is used to implement any of the methods described above by running the programs in the memory.
[0030] In a fourth aspect, an embodiment of the present application provides a sweeping machine, comprising: a controller and an information collector;
[0031] The information collector is used to collect visual images of specific targets;
[0032] The controller is used to determine the scene in which the sweeping robot is located based on the visual image of the specific target collected by the information collector by executing any of the above methods.
[0033] In a fifth aspect, an embodiment of the present application provides a computer program product, the computer program product comprising a computer program, and when the computer program is executed by a processor, the computer program implements any of the above methods. Optionally, the computer program can be stored in a readable storage medium of a computer device or in the cloud; the processor of the computer device reads the computer program from the readable storage medium or the cloud.
[0034] The scene distinguishing method proposed in the present application can determine the location information of a specific target based on the visual image of the specific target collected by the mobile device, wherein the specific target includes a target with a fixed position, and then use the location information of the specific target to match the specific target with the scene map to determine the scene where the specific target is located, wherein the scene map is generated based on the location information of the specific target in each scene, and finally determine the scene where the mobile device is located based on the scene where the specific target is located. With such a setting, the specific target in the current scene can be determined based on the visual image, and then the scene can be carefully distinguished based on the location information of the specific target to determine the scene where the mobile device is located, so as to avoid the positioning result jumping back and forth between similar scenes and ensure that the overall function of the mobile device is not affected. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are merely embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0036] Figure 1 It is a schematic diagram of a feasible application scenario of the method for distinguishing scenarios provided in an embodiment of the present application.
[0037] Figure 2 This is another feasible application scenario diagram of the scene distinction method provided in the embodiment of the present application.
[0038] Figure 3 It is a flowchart of a method for distinguishing scenes provided in an embodiment of the present application.
[0039] Figure 4 It is a schematic diagram of converting a visual image into a depth map provided by an embodiment of the present application.
[0040] Figure 5 It is a structural diagram of a scene distinguishing device provided in an embodiment of the present application.
[0041] Figure 6 It is a structural schematic diagram of an electronic device provided in an embodiment of the present application.
[0042] Figure 7 It is a structural schematic diagram of a sweeping machine provided in an embodiment of the present application. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0044] Mobile devices need to determine the correct posture based on their location in the scene to ensure that they can choose the best route to reach the target location and complete the scheduled task. For example, by determining the correct posture, a sweeping robot can ensure that all areas that need to be cleaned are covered without repeated cleaning or missing certain places, and the robotic arm can accurately locate and grasp objects, etc.
[0045] Existing technical solutions are generally based on single-line radar for mapping and positioning. Single-line radar can only provide distance measurement in the horizontal plane at a specific height, which means it cannot perceive feature differences in the vertical direction. For similar scenes with the same horizontal layout but different heights, single-line radar cannot effectively distinguish them, causing the positioning results to jump back and forth between similar scenes, affecting the overall function of the mobile device. Therefore, how to distinguish similar scenes has become a technical problem that technicians in this field need to solve urgently.
[0046] Based on this, the present application proposes a scene distinguishing method, device, equipment, sweeper and computer program product. The technical solution determines the specific target in the current scene based on the visual image, and then distinguishes the scene in detail according to the location information of the specific target, thereby achieving the effect of distinguishing similar scenes.
[0047] Figure 1 A feasible application scenario of the scene distinction method is shown, such as Figure 1 In the scenario shown, a mobile device is provided.
[0048] The mobile device provided in the embodiment of the present application can be implemented as any terminal with mobile function and computing function, such as intelligent robot, intelligent home appliance and intelligent vehicle-mounted device. Among them, the intelligent robot can be a sweeping robot.
[0049] In the above-mentioned feasible application scenarios, the mobile device is able to collect visual images of specific targets, wherein the specific targets include targets with fixed positions, and then determine the location information of the specific targets based on the visual images of the specific targets. Using the location information of the specific targets, the specific targets are matched with the scene map to determine the scene where the specific targets are located, wherein the scene map is generated based on the location information of the specific targets in each scene. Finally, based on the scene where the specific targets are located, the scene where the mobile device is located is determined.
[0050] If, based on the scene in which the specific target is located, it is determined that the scene in which the mobile device is located includes at least two places, the mobile device can determine the three-dimensional information of the specific target based on the visual image of the specific target, and then use the three-dimensional information of the specific target to match the specific target with the three-dimensional map to determine the scene in which the specific target is located, wherein the three-dimensional map is generated based on the three-dimensional information of the specific target in each scene, and finally, based on the scene in which the specific target is located, the scene in which the mobile device is located is determined.
[0051] Figure 2 Another feasible application scenario of the scene distinction method is shown, such as Figure 2 In the scenario shown, a mobile device and a server are provided.
[0052] The mobile device provided in the embodiment of the present application can be implemented as any terminal with mobile functions, such as an intelligent robot, an intelligent home appliance, and an intelligent vehicle-mounted device. Among them, the intelligent robot can be a sweeping robot. The server provided in the embodiment of the present application can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks (CDN, Content Delivery Network), and big data and artificial intelligence platforms.
[0053] The mobile device and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in the embodiments of the present application. For example, the mobile device and the server can communicate through the target network, which can be a network that can be subdivided into multiple subnetworks, and the target network or the multiple subnetworks contained in the target network can be at least one of a cellular mobile network (such as 2G, 3G, 4G or 5G), ZIGBEE, Wi-Fi, and Bluetooth, or any combination of at least one of these networks and other networks.
[0054] In the above-mentioned feasible application scenarios, the mobile device is able to collect visual images of specific targets, wherein the specific targets include targets with fixed positions, and then upload the visual images of the specific targets to the server. The server determines the location information of the specific targets based on the visual images of the specific targets collected by the mobile device, and then uses the location information of the specific targets to match the specific targets with the scene map to determine the scene where the specific targets are located, wherein the scene map is generated based on the location information of the specific targets in each scene, and finally, based on the scene where the specific targets are located, the scene where the mobile device is located is determined.
[0055] If it is determined that the scene where the mobile device is located includes at least two places based on the scene where the specific target is located, the server can determine the three-dimensional information of the specific target based on the visual image of the specific target, and then use the three-dimensional information of the specific target to match the specific target with the three-dimensional map to determine the scene where the specific target is located, wherein the three-dimensional map is generated based on the three-dimensional information of the specific target in each scene, and finally, based on the scene where the specific target is located, the scene where the mobile device is located is determined.
[0056] With this setting, the specific target in the current scene can be determined based on the visual image, and then the scene can be carefully distinguished according to the location information of the specific target to determine the scene where the mobile device is located, thereby avoiding the positioning result jumping back and forth between similar scenes and ensuring that the overall function of the mobile device is not affected.
[0057] An exemplary embodiment of the present specification provides a method for distinguishing scenes, and the method is described by taking an electronic device as an example. The electronic device may be Figure 1 The mobile device or Figure 2 The server in the illustrated embodiment. Figure 3 As shown, the method includes:
[0058] S101. Determine location information of a specific target based on a visual image of the specific target collected by a mobile device.
[0059] The above-mentioned specific target refers to a target with a fixed position. In other words, the position of the specific target is relatively fixed and will not be easily moved. For example, a bed or a dining table in a room. In some embodiments, the specific target can be set according to actual conditions, which is not limited in this embodiment.
[0060] A visual image refers to a representation of information captured by an optical device that can be perceived and interpreted by the human eye or a machine vision system. A visual image can be a static photo or a dynamic video frame, which records the light intensity, color, texture and other characteristics of the scene and contains rich spatial and temporal information. In some embodiments, the visual image is an RGB image.
[0061] The mobile device of the present application is provided with an optical device for collecting visual images of the scene where the mobile device is located. The optical device can be installed around the mobile device to obtain visual images of the scene where the mobile device is located. The optical device can be a camera, etc., which is not limited in this embodiment.
[0062] When the acquired visual image of the scene where the mobile device is located contains a specific target, the visual image of the scene is determined to be a visual image of the specific target. Specifically, the target detection technology can be used to detect whether the visual image of the scene contains the specific target. When it is detected that the visual image of the scene contains the specific target, the visual image of the scene is determined to be a visual image of the specific target.
[0063] The position information of the specific target is determined by analyzing the visual image of the specific target. In some embodiments, the image pixel coordinates of the specific target can be determined based on the visual image of the specific target, and then the image pixel coordinates of the specific target can be converted to the world coordinate system in combination with the camera internal parameters to obtain the coordinates of the specific target in the world coordinate system as the position information of the specific target.
[0064] S102: Using the location information of the specific target, match the specific target with the scene map to determine the scene where the specific target is located.
[0065] The scene map refers to a map of each scene in the working location of the mobile device. For example, if the mobile device is a sweeping robot, and the working location of the sweeping robot is inside a house, then each room in the house corresponds to a scene, and the scene map includes maps of each room.
[0066] The scene map is generated based on the location information of specific targets in each scene. Specifically, when the scene map is not constructed, the working place of the mobile device is an unknown environment for the mobile device, and the mobile device needs to move in the working place, determine its own position during the movement, and create the scene map. In some embodiments, the Simultaneous Localization And Mapping (SLAM) technology is used to control the movement of the mobile device in the working place, determine its own position during the movement, and create the scene map.
[0067] More specifically, during the movement of the mobile device, it acquires a visual image of the scene in which it is located. When the acquired visual image of the scene in which it is located contains a specific target, the visual image is determined to be the visual image of the specific target. The location information of the specific target is determined by analyzing the visual image of the specific target, and then the location information of the specific target is refreshed into the scene map. It should be noted that when generating the scene map and determining the scene in which the specific target is located in the above steps, it is necessary to analyze the visual image of the specific target to determine the location information of the specific target. In order to ensure the consistency of the two and facilitate position matching in subsequent steps, when generating the scene map and determining the scene in which the specific target is located in the above steps, the same method is used to analyze the visual image of the specific target to determine the location information of the specific target.
[0068] For example, if the mobile device is a sweeping robot, the specific targets include sofas, beds, dining tables and cabinets, and the working place of the sweeping robot is inside the house, the sweeping robot is first controlled to move inside the house to obtain a visual image of the scene in which it is located. When the obtained visual image of the scene in which it is located includes a sofa, it is determined that the visual image is a visual image of the sofa. By analyzing the visual image of the sofa, the location information of the sofa is determined, and then the location information of the sofa is refreshed to the scene map. In the same way, the location information of specific targets such as the bed, dining table and cabinet is refreshed to the scene map to obtain a created scene map.
[0069] In some embodiments, the scene map may also be divided into scenes, for example, the scene map may be divided into a bedroom, a living room, and a dining room, etc., which is not limited in this embodiment.
[0070] The specific target is matched with the scene map using the location information of the specific target to determine the specific scene of the multiple scenes included in the scene map. In some embodiments, the scene where the location information of the specific target is located in the scene map can be determined as the scene where the specific target is located based on the location information of the specific target. In this way, the scene can be finely distinguished by the location information of the specific target to determine the scene where the specific target is located.
[0071] S103: Determine the scene where the mobile device is located according to the scene where the specific target is located.
[0072] In some embodiments, the scene where the specific target is located is determined to be the scene where the mobile device is located. For example, if the scene where the specific target is located is determined to be a living room, it can be determined that the scene where the mobile device is located is also a living room.
[0073] In the above embodiments, the location information of a specific target can be determined based on the visual image of the specific target collected by the mobile device, wherein the specific target includes a target with a fixed position, and then the location information of the specific target is used to match the specific target with the scene map to determine the scene in which the specific target is located, wherein the scene map is generated based on the location information of the specific target in each scene, and finally the scene in which the mobile device is located is determined based on the scene in which the specific target is located. With such a setting, the specific target in the current scene can be determined based on the visual image, and then the scene can be carefully distinguished based on the location information of the specific target to determine the scene in which the mobile device is located, thereby avoiding the positioning result jumping back and forth between similar scenes and ensuring that the overall function of the mobile device is not affected.
[0074] As an optional implementation, another embodiment of the present application discloses that the steps of the above embodiment are based on the visual image of the specific target collected by the mobile device to determine the location information of the specific target, which may specifically include the following steps:
[0075] Based on the visual image of the specific target collected by the mobile device, the coordinates of the grounding line of the specific target are determined as the position information of the specific target.
[0076] The grounding line of the specific target includes a line formed when the specific target contacts the ground. In an embodiment of the present application, the coordinates of the grounding line of the specific target are calculated as the position information of the specific target. In this way, by only calculating the coordinates of the grounding line of the specific target, the amount of calculation can be effectively reduced and the speed of distinguishing scenes can be improved.
[0077] As an optional implementation, another embodiment of the present application discloses that the steps of the above embodiment are based on the visual image of the specific target collected by the mobile device to determine the coordinates of the ground line of the specific target as the location information of the specific target, which may specifically include the following steps:
[0078] Based on the visual image of the specific target captured by the mobile device, the image pixel coordinates of the grounding line of the specific target are determined; the image pixel coordinates of the grounding line of the specific target are subjected to an inverse perspective transformation to obtain the coordinates of the grounding line of the specific target in a bird's-eye view, and the coordinates of the grounding line of the specific target in a bird's-eye view are determined as the position information of the specific target.
[0079] Specifically, the visual image of the specific target can be used to determine the image pixel coordinates of the grounding line of the specific target, and then the image pixel coordinates of the grounding line of the specific target can be inversely transformed. Specifically, the image pixel coordinates of the grounding line of the specific target can be inversely transformed according to the inverse perspective transformation matrix to obtain the coordinates of the grounding line of the specific target in a bird's-eye view, so as to determine the coordinates of the grounding line of the specific target in a bird's-eye view as the location information of the specific target.
[0080] The inverse perspective transformation matrix can be obtained by the following steps:
[0081] For a point A in the world coordinate system, it can be converted into image pixel coordinates by the following formula:
[0082]
[0083] Among them, k cam is the camera internal parameter, is the transformation matrix from the world coordinate system to the camera coordinate system, z cam is the normalization coefficient, is the coordinate of A in the world coordinate system, is the image pixel coordinate of A.
[0084] Furthermore, assuming that the ground is flat, the world coordinates (X, Y, 0) can be mapped to the top view to obtain the coordinates (u, v). That is, the world coordinates (X, Y, 0) and the top view coordinates (u, v) can be obtained by calculation or measurement. The world coordinates (X, Y, 0) and the top view coordinates (u, v) of the same point can be used as a point pair. In this embodiment, such point pairs are used to solve the inverse perspective transformation matrix M, where the perspective transformation matrix M can be a 3×3 matrix. A set of linear equations can be constructed for solution, with at least four corresponding point pairs, and no three or more points can be collinear. The formula is as follows:
[0085]
[0086] Based on the above formula, the inverse perspective transformation matrix M can be obtained.
[0087] Based on the image pixel coordinates of the grounding line of the specific target determined by the embodiment of the present application and the inverse perspective transformation matrix, the coordinate a of the grounding line of the specific target in the top view can be obtained. The specific formula is as follows:
[0088]
[0089] Where k represents the image pixel coordinate of the ground line of a specific target, and M represents the inverse perspective transformation matrix.
[0090] It should be noted that when generating a scene map and determining the scene in which a specific target is located, it is necessary to analyze the visual image of the specific target to determine the location information of the specific target. In order to ensure the consistency of the two and facilitate position matching in subsequent steps, when generating a scene map and determining the scene in which a specific target is located in the above steps, the same method is used to analyze the visual image of the specific target to determine the location information of the specific target.
[0091] In the above embodiment, the coordinates of the grounding line of a specific target in a top-down perspective can be calculated. Converting the grounding line of a specific target to a top-down perspective makes objects and structures in the plane more intuitive, facilitates obstacle identification and cleaning path planning, and eliminates the complexity caused by perspective deformation and reduces the amount of calculation.
[0092] As an optional implementation, another embodiment of the present application discloses that the steps of the above embodiment determine the scene where the mobile device is located according to the scene where the specific target is located, and specifically may include the following steps:
[0093] The point cloud map and the radar point cloud collected by the mobile device are matched to obtain a scene matching result; the point cloud map is generated based on the radar point cloud in each scene; based on the scene matching result and the scene where the specific target is located, the scene where the mobile device is located is determined.
[0094] In the embodiment of the present application, the scene where the mobile device is located is determined from two dimensions: radar point cloud and visual image.
[0095] The mobile device can collect the radar point cloud of the scene it is in, and then match the point cloud map with the radar point cloud collected by the mobile device to obtain the scene matching result. Among them, the point cloud map is generated based on the radar point cloud in each scene. The mobile device can use a single-line radar to collect the radar point cloud of the scene it is in, and then use SLAM technology to determine its own position while creating an environmental point cloud map.
[0096] Specifically, a single-line radar will generate a circle of radar point clouds when it hits an obstacle. Based on the matching of this circle of radar point clouds and the point cloud map, the existing matching algorithms, such as Cartographer, can be used to obtain the position and posture of the mobile device. The current position and posture are obtained based on the current map, and then the current point cloud is refreshed to the map based on the current position and posture to achieve map update.
[0097] In this embodiment, the point cloud map and the radar point cloud collected by the mobile device are matched to obtain a scene matching result, that is, the scene where the mobile device is located determined based on the radar point cloud. Then, the scene where the mobile device is located is determined by combining the scene matching result and the scene where the specific target is located.
[0098] In some embodiments, the point cloud map and the radar point cloud collected by the mobile device can be matched to obtain the scene where the mobile device is located according to the radar point cloud, and the specific target can be matched with the scene map to determine the scene where the specific target is located. If the scene where the mobile device is located according to the radar point cloud is the same as the scene where the specific target is located, then the scene is determined to be the scene where the mobile device is located; if the scene where the mobile device is located according to the radar point cloud is different from the scene where the specific target is located, then the radar point cloud and the visual image of the specific target can be re-collected according to the steps of the above embodiment to distinguish the scenes again; if the scene where the mobile device is located according to the radar point cloud includes multiple scenes, then the scene where the specific target is located is selected from the above multiple scenes as the scene where the mobile device is located.
[0099] In the above embodiment, the scene in which the mobile device is located can be determined from two dimensions: radar point cloud and visual image. This allows for detailed distinction of scenes, avoids positioning results jumping back and forth between similar scenes, and ensures that the overall functionality of the mobile device is not affected.
[0100] As an optional implementation, another embodiment of the present application discloses that if, according to the scene in which the specific target is located, it is determined that the scene in which the mobile device is located includes at least two places, the method of the above embodiment further includes:
[0101] Based on the visual image of the specific target collected by the mobile device, the three-dimensional information of the specific target is determined; using the three-dimensional information of the specific target, the specific target is matched with the three-dimensional map to determine the scene where the specific target is located; the three-dimensional map is generated based on the three-dimensional information of the specific target in each scene; based on the scene where the specific target is located, the scene where the mobile device is located is determined.
[0102] In some embodiments, if the similarity between scenes is particularly high and effective scene differentiation is still not possible from the two dimensions of radar point cloud and visual image, the three-dimensional information of the specific target can be determined based on the visual image of the specific target collected by the mobile device.
[0103] Specifically, the depth information of the specific target can be determined based on the visual image of the specific target captured by the mobile device, and the coordinates of the specific target in the three-dimensional space can be determined as the three-dimensional information of the specific target using the depth information of the specific target and the coordinates of the mobile device in the three-dimensional space.
[0104] In some embodiments, the depth information of a specific target can be extracted based on a deep learning network. Using a dataset with depth annotations, the dataset is input into a deep learning network, the output depth map is calculated by forward propagation, and then the difference between the predicted depth map and the true depth map is calculated according to the loss function, and the parameters of the deep learning network are adjusted by back propagation. Continuously iterate the training so that the deep learning network gradually learns the mapping relationship from image to depth map. The deep learning network can be a convolutional neural network or an encoder-decoder structure, which is not limited in this embodiment.
[0105] Input the visual image of a specific target into the trained deep learning network to obtain the depth map output by the deep learning network, such as Figure 4 As shown. By combining the depth map of the specific target and the parameters of the camera, the distance between the specific target and the camera, and the angle between the line connecting the specific target and the camera and the horizontal plane can be calculated, and the distance between the specific target and the camera and the angle between the line connecting the specific target and the camera and the horizontal plane are used as the depth information of the specific target.
[0106] It should be noted that the visual image of the above-mentioned specific target can be an RGB image taken by a single camera.
[0107] Furthermore, based on the coordinates of the mobile device in the three-dimensional space and the depth information of the specific target, the coordinates of the specific target in the three-dimensional space can be calculated, and the coordinates of the specific target in the three-dimensional space can be determined as the three-dimensional information of the specific target.
[0108] The scene map refers to a map of each scene in the working location of the mobile device. For example, if the mobile device is a sweeping robot, and the working location of the sweeping robot is inside a house, then each room in the house corresponds to a scene, and the scene map includes maps of each room.
[0109] The three-dimensional map is generated based on the three-dimensional information of specific targets in each scene. Specifically, the mobile device needs to move in the work place, determine its own position during the movement, and create a three-dimensional map. In some embodiments, SLAM technology is used to control the mobile device to move in the work place, determine its own position during the movement, and create a three-dimensional map.
[0110] More specifically, the mobile device acquires a visual image of the scene in which it is located during movement, and when the acquired visual image of the scene in which it is located contains a specific target, the visual image is determined to be a visual image of the specific target. The visual image of the specific target is analyzed according to the description of the above embodiment to determine the three-dimensional information of the specific target, and then the three-dimensional information of the specific target is updated to the three-dimensional map.
[0111] For example, if the mobile device is a sweeping robot, the specific targets include sofas, beds, dining tables and cabinets, and the working place of the sweeping robot is inside a house, the sweeping robot is first controlled to move inside the house to obtain a visual image of the scene in which it is located. When the obtained visual image of the scene in which it is located includes a sofa, it is determined that the visual image is a visual image of the sofa. By analyzing the visual image of the sofa, the three-dimensional information of the sofa is determined, and then the three-dimensional information of the sofa is refreshed into the three-dimensional map. In the same way, the three-dimensional information of specific targets such as the bed, dining table and cabinet is refreshed into the three-dimensional map to obtain a created three-dimensional map.
[0112] In some embodiments, the three-dimensional map may also be divided into scenes, for example, the three-dimensional map may be divided into a bedroom, a living room, and a dining room, etc., which is not limited in this embodiment.
[0113] The three-dimensional information of the specific target is used to match the specific target with the three-dimensional map, and determine the specific scene of the specific target in the multiple scenes included in the three-dimensional map. In some embodiments, the scene where the three-dimensional information of the specific target is located in the three-dimensional map can be determined as the scene where the specific target is located. In this way, the three-dimensional information of the specific target can be used to make a detailed distinction between the scenes and determine the scene where the specific target is located. Then, the scene where the specific target is located is determined to be the scene where the mobile device is located.
[0114] In the above embodiment, when it is determined that the scene where the mobile device is located includes at least two places according to the scene where the specific target is located, further matching is performed according to the three-dimensional information of the specific target to determine the scene where the specific target is located.
[0115] Corresponding to the above-mentioned scene distinguishing method, the present application embodiment also discloses a scene distinguishing device, see Figure 5 As shown, the device comprises:
[0116] The first determination unit 100 is used to determine the location information of the specific target based on the visual image of the specific target collected by the mobile device; the specific target includes a target with a fixed position;
[0117] The matching unit 110 is used to match the specific target with the scene map using the location information of the specific target to determine the scene where the specific target is located; the scene map is generated based on the location information of the specific target in each scene;
[0118] The second determining unit 120 is used to determine the scene where the mobile device is located according to the scene where the specific target is located.
[0119] As an optional implementation, another embodiment of the present application discloses that the first determining unit 110 in the above embodiment is specifically configured to:
[0120] Based on the visual image of the specific target collected by the mobile device, the coordinates of the grounding line of the specific target are determined as the position information of the specific target; the grounding line of the specific target includes a line formed when the specific target contacts the ground.
[0121] As an optional implementation, another embodiment of the present application discloses that the first determining unit 110 in the above embodiment is specifically configured to:
[0122] Based on the visual image of the specific target captured by the mobile device, the image pixel coordinates of the grounding line of the specific target are determined; the image pixel coordinates of the grounding line of the specific target are subjected to an inverse perspective transformation to obtain the coordinates of the grounding line of the specific target in a bird's-eye view, and the coordinates of the grounding line of the specific target in a bird's-eye view are determined as the position information of the specific target.
[0123] As an optional implementation, another embodiment of the present application discloses that the second determining unit 120 in the above embodiment is specifically configured to:
[0124] The point cloud map and the radar point cloud collected by the mobile device are matched to obtain a scene matching result; the point cloud map is generated based on the radar point cloud in each scene; based on the scene matching result and the scene where the specific target is located, the scene where the mobile device is located is determined.
[0125] As an optional implementation, another embodiment of the present application discloses that the device in the above embodiment further includes:
[0126] The third determination unit is used to determine the three-dimensional information of the specific target based on the visual image of the specific target collected by the mobile device; use the three-dimensional information of the specific target to match the specific target with the three-dimensional map to determine the scene where the specific target is located; the three-dimensional map is generated based on the three-dimensional information of the specific target in each scene; and determine the scene where the mobile device is located based on the scene where the specific target is located.
[0127] As an optional implementation method, another embodiment of the present application discloses that the third determination unit of the above embodiment is specifically used to: determine the depth information of the specific target based on the visual image of the specific target collected by the mobile device; use the depth information of the specific target and the coordinates of the mobile device in the three-dimensional space to determine the coordinates of the specific target in the three-dimensional space as the three-dimensional information of the specific target.
[0128] Specifically, the device provided in this embodiment belongs to the same application concept as the method provided in the above embodiment of this application, can execute the method provided in any of the above embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For the technical details not fully described in this embodiment, please refer to the specific processing content of the method provided in the above embodiment of this application, which will not be repeated here.
[0129] The functions implemented by the above units can be implemented by the same or different processors respectively, and the embodiments of the present application are not limited thereto.
[0130] It should be understood that the units in the above devices can be implemented in the form of a processor calling software. For example, the device includes a processor, the processor is connected to a memory, and instructions are stored in the memory. The processor calls the instructions stored in the memory to implement any of the above methods or realize the functions of each unit of the device, wherein the processor can be a general-purpose processor, such as a CPU or a microprocessor, etc., and the memory can be a memory in the device or a memory outside the device. Alternatively, the units in the device can be implemented in the form of hardware circuits, and the functions of some or all units can be realized by designing the hardware circuits. The hardware circuit can be understood as one or more processors; for example, in one implementation, the hardware circuit is an ASIC, and the functions of some or all of the above units are realized by designing the logical relationship of the components in the circuit; for another example, in another implementation, the hardware circuit can be implemented by PLD, taking FPGA as an example, which can include a large number of logic gate circuits, and the connection relationship between the logic gate circuits is configured by the configuration file, so as to realize the functions of some or all of the above units. All units of the above devices can be implemented in the form of a processor calling software, or in the form of hardware circuits, or in part by a processor calling software, and the remaining part is implemented in the form of hardware circuits.
[0131] In an embodiment of the present application, a processor is a circuit with the ability to process signals. In one implementation, the processor may be a circuit with the ability to read and run instructions, such as a CPU, a microprocessor, a GPU, or a DSP; in another implementation, the processor may implement certain functions through the logical relationship of a hardware circuit, and the logical relationship of the hardware circuit is fixed or reconfigurable, such as a hardware circuit implemented by an ASIC or PLD, such as an FPGA. In a reconfigurable hardware circuit, the process of the processor loading a configuration document to implement the hardware circuit configuration can be understood as the process of the processor loading instructions to implement the functions of some or all of the above units. In addition, it can also be a hardware circuit designed for artificial intelligence, which can be understood as an ASIC, such as an NPU, TPU, DPU, etc.
[0132] It can be seen that each unit in the above device can be one or more processors (or processing circuits) configured to implement the above method, such as: CPU, GPU, NPU, TPU, DPU, microprocessor, DSP, ASIC, FPGA, or a combination of at least two of these processor forms.
[0133] In addition, all or part of the units in the above device can be integrated together, or can be implemented independently. In one implementation, these units are integrated together and implemented in the form of a SOC. The SOC may include at least one processor for implementing any of the above methods or implementing the functions of each unit of the device. The type of the at least one processor may be different, for example, including a CPU and an FPGA, a CPU and an artificial intelligence processor, a CPU and a GPU, etc.
[0134] An embodiment of the present application further provides a control device, which includes a processor and an interface circuit. The processor in the control device is connected to an input-output component through the interface circuit of the control device.
[0135] The input-output component specifically refers to a hardware component that enables a user to input information and output information to the user, such as a microphone, keyboard, handwriting tablet, touch screen, display, speaker, printer, etc.
[0136] The above-mentioned interface circuit can be any interface circuit that can realize the data communication function, for example, it can be a USB interface circuit, a Type-C interface circuit, a serial port circuit, a PCIE circuit, etc.
[0137] The processor in the control device is a circuit with signal processing capability, which distinguishes scenes in detail by executing any of the scene distinguishing methods introduced in the above embodiments. The specific implementation of the processor can refer to the above processor implementation, and the embodiments of this application are not strictly limited.
[0138] When the control device is applied to a device with a human-computer interaction function, the input and output components of the control device may be input components and output components on the device, such as a microphone, a keyboard, a handwriting tablet, a touch screen, a display, an audio player, etc. At the same time, the processor of the control device may be a CPU or GPU, etc. provided by the device, and the interface circuit of the control device may be an interface circuit between the information input component of the device and a processor such as a CPU or GPU.
[0139] Corresponding to the above-mentioned scene distinguishing method, the present application embodiment also discloses an electronic device, see Figure 6 As shown, the electronic device includes:
[0140] Memory 200 and processor 210;
[0141] The memory 200 is connected to the processor 210 and is used to store programs;
[0142] The processor 210 is configured to implement the scene distinguishing method disclosed in any of the above embodiments by running the program stored in the memory 200 .
[0143] Specifically, the electronic device may further include: a bus, a communication interface 220 , an input device 230 and an output device 240 .
[0144] The processor 210, the memory 200, the communication interface 220, the input device 230 and the output device 240 are connected to each other via a bus.
[0145] A bus may include a pathway that transfers information between components of a computer system.
[0146] The processor 210 may be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits for controlling the execution of the program of the present application. It may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0147] The processor 210 may include a main processor, and may also include a baseband chip, a modem, and the like.
[0148] The memory 200 stores a program for executing the technical solution of the present application, and may also store an operating system and other key services. Specifically, the program may include a program code, and the program code includes computer operation instructions. More specifically, the memory 200 may include a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM), other types of dynamic storage devices that can store information and instructions, a disk storage, a flash, and the like.
[0149] The input device 230 may include a device for receiving data and information input by a user, such as a keyboard, a mouse, a camera, a scanner, a light pen, a voice input device, a touch screen, a pedometer, or a gravity sensor.
[0150] Output device 240 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0151] The communication interface 220 may include any transceiver or the like to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0152] The processor 210 executes the program stored in the memory 200 and calls other devices, which can be used to implement the various steps of the scene distinguishing method provided in the above embodiments of the present application.
[0153] Another embodiment of the present application also provides a sweeping machine, see Figure 7 As shown, the sweeping machine includes a controller 300 and an information collector 310. The information collector 310 is used to collect visual images of specific targets. The controller 300 is used to determine the scene in which the sweeping machine is located based on the visual images of the specific targets collected by the information collector 310 by executing the scene distinguishing method provided in the above-mentioned embodiment of the present application.
[0154] The sweeping machine provided in this embodiment belongs to the same application concept as the scene distinguishing method provided in the above-mentioned embodiments of this application, can execute the scene distinguishing method provided in any of the above-mentioned embodiments of this application, and has functional modules and beneficial effects corresponding to executing the above-mentioned scene distinguishing method. For technical details not fully described in this embodiment, please refer to the specific processing content of the scene distinguishing method provided in the above-mentioned embodiments of this application, which will not be repeated here.
[0155] In addition to the above methods and devices, the embodiments of the present application may also be computer program products, which include computer programs. When the computer programs are executed by a processor, the scene differentiation method provided by any of the above embodiments of the present application may be executed. Optionally, the computer program may be stored in a readable storage medium or in the cloud of a computer device; the processor of the computer device reads the computer program from the readable storage medium or the cloud.
[0156] The computer program product may be written in any combination of one or more programming languages to write program codes for performing the operations of the embodiments of the present application, including object-oriented programming languages such as Java, C++, etc., and also conventional procedural programming languages such as "C" language or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0157] The computer program product may be implemented in hardware, software or a combination thereof. In one optional embodiment, the computer program product is embodied as a computer storage medium, and in another optional embodiment, the computer program product is embodied as a software product, such as a software development kit (SDK).
[0158] In addition, the embodiments of the present application may also be a computer-readable storage medium on which computer program instructions are stored. When the computer program instructions are executed by a processor, the processor executes each step of the scene distinguishing method provided in the above embodiments.
[0159] The computer readable storage medium may adopt any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may include, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples (non-exhaustive list) of readable storage media include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0160] Specifically, the specific working contents of each part of the above-mentioned electronic device, computer program product and storage medium, as well as the specific processing contents when the computer program product or the computer program on the above-mentioned storage medium is executed by the processor, can all be found in the contents of the various embodiments of the above-mentioned scene distinguishing method, and will not be repeated here.
[0161] For the aforementioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the order of the actions described, because according to the present application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required for the present application.
[0162] It should be noted that each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments, and the same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0163] The steps in the methods of each embodiment of the present application can be adjusted in order, combined and deleted according to actual needs, and the technical features recorded in each embodiment can be replaced or combined.
[0164] The modules and sub-modules in the devices and terminals in the various embodiments of the present application can be combined, divided and deleted according to actual needs.
[0165] In the several embodiments provided in the present application, it should be understood that the disclosed terminals, devices and methods can be implemented in other ways. For example, the terminal embodiments described above are only schematic, for example, the division of modules or submodules is only a logical function division, and there may be other division methods in actual implementation, such as multiple submodules or modules can be combined or integrated into another module, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0166] The modules or submodules described as separate components may or may not be physically separated, and the components of the modules or submodules may or may not be physical modules or submodules, that is, they may be located in one place, or they may be distributed on multiple network modules or submodules. Some or all of the modules or submodules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0167] In addition, each functional module or submodule in each embodiment of the present application may be integrated into one processing module, or each module or submodule may exist physically separately, or two or more modules or submodules may be integrated into one module. The above-mentioned integrated modules or submodules may be implemented in the form of hardware or in the form of software functional modules or submodules.
[0168] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0169] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly by hardware, software units executed by a processor, or a combination of the two. The software units may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0170] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "including a..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0171] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to the embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for distinguishing scenes, characterized in that: include: Determining location information of the specific target based on a visual image of the specific target captured by the mobile device; The specific targets include targets with fixed positions; Using the location information of the specific target, matching the specific target with a scene map is performed to determine the scene where the specific target is located; The scene map is generated based on the location information of specific targets in each scene; The scene where the mobile device is located is determined according to the scene where the specific target is located.
2. The method according to claim 1, characterized in that The determining the location information of the specific target based on the visual image of the specific target collected by the mobile device includes: Based on a visual image of a specific target captured by a mobile device, coordinates of a grounding line of the specific target are determined as position information of the specific target; the grounding line of the specific target includes a line formed when the specific target contacts the ground.
3. The method according to claim 2, characterized in that The step of determining the coordinates of the grounding line of the specific target as the position information of the specific target based on the visual image of the specific target collected by the mobile device includes: Determining image pixel coordinates of a ground line of the specific target based on a visual image of the specific target captured by the mobile device; Perform an inverse perspective transformation on the image pixel coordinates of the grounding line of the specific target to obtain the coordinates of the grounding line of the specific target in a bird's-eye view, and determine the coordinates of the grounding line of the specific target in a bird's-eye view as the position information of the specific target.
4. The method according to claim 1, characterized in that: The determining the scene where the mobile device is located according to the scene where the specific target is located includes: Matching the point cloud map with the radar point cloud collected by the mobile device to obtain a scene matching result; the point cloud map is generated based on the radar point cloud in each scene; The scene where the mobile device is located is determined according to the scene matching result and the scene where the specific target is located.
5. The method according to claim 1, characterized in that If it is determined that the scene where the mobile device is located includes at least two scenes according to the scene where the specific target is located, the method further includes: Determining three-dimensional information of the specific target based on a visual image of the specific target acquired by the mobile device; Using the three-dimensional information of the specific target, matching the specific target with the three-dimensional map to determine the scene where the specific target is located; the three-dimensional map is generated based on the three-dimensional information of the specific target in each scene; The scene where the mobile device is located is determined according to the scene where the specific target is located.
6. The method according to claim 5, characterized in that The determining of the three-dimensional information of the specific target based on the visual image of the specific target acquired by the mobile device includes: Determining depth information of a specific target based on a visual image of the specific target captured by a mobile device; The coordinates of the specific target in the three-dimensional space are determined as the three-dimensional information of the specific target by using the depth information of the specific target and the coordinates of the mobile device in the three-dimensional space.
7. A scene distinguishing device, characterized in that: include: A first determining unit, configured to determine location information of a specific target based on a visual image of the specific target acquired by a mobile device; The specific targets include targets with fixed positions; A matching unit, used to match the specific target with a scene map by using the location information of the specific target, and determine the scene where the specific target is located; The scene map is generated based on the location information of specific targets in each scene; The second determining unit is used to determine the scene where the mobile device is located according to the scene where the specific target is located.
8. An electronic device, characterized in that: include: Memory and processor; Wherein, the memory is used to store programs; The processor is used to implement the method according to any one of claims 1 to 6 by running the program in the memory.
9. A sweeping machine, characterized in that: include: Controller and information collector; The information collector is used to collect visual images of specific targets; The controller is used to determine the scene in which the sweeping robot is located based on the visual image of the specific target collected by the information collector by executing the method described in any one of claims 1 to 6.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.