Semantic annotation method, device and system
By performing scene reconstruction and semantic annotation on three-dimensional scenes, the problem of low accuracy of three-dimensional scene annotation in traditional technology is solved, and high-accuracy three-dimensional semantic annotation is achieved.
Patent Information
- Application Number
- CN202010724336.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-07-24
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2040-07-24
AI Technical Summary
The traditional semantic annotation method has low accuracy in three-dimensional scenes, making it difficult to effectively support semantic annotation of three-dimensional scenes.
By obtaining the scene video sequence of the three-dimensional scene sent by the front end, scene reconstruction is carried out, and description information of the reconstructed three-dimensional scene is obtained, and sent to the front end to obtain semantic annotation results.
The semantic annotation accuracy of three-dimensional scenes is improved, and accurate reconstruction and semantic annotation of three-dimensional scenes are realized.
Smart Images

Figure CN111860370B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision technology, and in particular, to a semantic annotation method, apparatus, and system. Background Art
[0002] Semantic annotation is used to solve the problem of which target each point in a scene belongs to. For example, for an indoor scene, semantic annotation is used to determine that the category to which each point in the scene belongs is a table, a chair, a computer, and so on. When traditional semantic annotation methods are used for annotating a three-dimensional scene, the annotation accuracy is relatively low. Summary of the Invention
[0003] The present disclosure provides a semantic annotation method, apparatus, and system.
[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a semantic annotation method applied to a server. The method includes: obtaining a scene video sequence of a three-dimensional scene sent by a front end; performing scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain description information of the reconstructed three-dimensional scene; and sending the description information of the three-dimensional scene to the front end to obtain a semantic annotation result of the three-dimensional scene returned by the front end, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene.
[0005] In some embodiments, each frame image in the scene video sequence includes an R-channel image, a G-channel image, a B-channel image, and a depth image of the three-dimensional scene.
[0006] In some embodiments, the performing scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain description information of the reconstructed three-dimensional scene includes: performing scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain a plurality of meshes corresponding to the three-dimensional scene; and respectively obtaining description information of each mesh in the plurality of meshes, where the description information of the three-dimensional scene includes the description information of each mesh.
[0007] In some embodiments, the method further includes: obtaining a semantic label of each point in the three-dimensional scene; after obtaining the plurality of meshes corresponding to the three-dimensional scene, generating a semantic label of each mesh according to the semantic labels of each point in each mesh in the plurality of meshes, where each mesh includes at least one point in the three-dimensional scene; and sending the semantic label of each mesh to the front end to obtain a correction result of the semantic annotation result obtained based on the semantic label of each mesh at the front end.
[0008] In some embodiments, the method further includes: after sending the description information of the three-dimensional scene to the front end, obtaining the semantic annotation result returned by the front end; projecting the semantic annotation result onto each frame image of the scene video sequence.
[0009] In some embodiments, the number of the front ends is multiple.
[0010] In some embodiments, the method further includes: respectively obtaining the semantic annotation results of each of the multiple front ends; saving the semantic annotation results of each front end according to the scenes corresponding to the semantic annotation results of each front end.
[0011] According to a second aspect of the embodiments of the present disclosure, a semantic annotation method is provided, which is applied to a front end. The method includes: sending a scene video sequence of a three-dimensional scene to a server, so that the server performs scene reconstruction on the three-dimensional scene according to the scene video sequence; receiving the description information of the three-dimensional scene returned by the server after the scene reconstruction; generating the three-dimensional scene according to the description information of the three-dimensional scene, and returning the semantic annotation result of the three-dimensional scene to the server, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene.
[0012] In some embodiments, the description information of the reconstructed three-dimensional scene includes the description information of each of multiple meshes in the three-dimensional scene; the semantic annotation result of the three-dimensional scene is obtained based on the following method: aggregating the multiple meshes according to the description information of each of the multiple meshes to obtain at least one aggregated mesh, and each aggregated mesh in the at least one aggregated mesh corresponds to an object in the three-dimensional scene; performing semantic annotation on each of the at least one aggregated meshes to obtain the semantic annotation result of the three-dimensional scene.
[0013] In some embodiments, the method further includes: hiding at least one first semantic annotation result in the semantic annotation result; and / or displaying at least one first semantic annotation result that has been hidden.
[0014] In some embodiments, the hiding of at least one first semantic annotation result in the semantic annotation result includes: generating a set of cut surfaces for each of the at least one first semantic annotation results, where a set of cut surfaces of a first semantic annotation result wraps the first semantic annotation result; hiding each of the at least one first semantic annotation results through the set of cut surfaces of each of the at least one first semantic annotation results.
[0015] In some embodiments, the semantic annotation of each aggregation grid in the at least one aggregation grid includes: receiving a selection instruction for each aggregation grid in the at least one aggregation grid; generating a bounding box for each aggregation grid according to the corresponding selection instruction of each aggregation grid, where the bounding box of an aggregation grid includes each point in the aggregation grid, and the bounding box is a bounding box or a convex hull; and performing semantic annotation on the points in the bounding box of each aggregation grid.
[0016] In some embodiments, the generation of the three-dimensional scene according to the description information of the three-dimensional scene is executed by a first thread; the semantic annotation result of the three-dimensional scene is obtained by a second thread; wherein, the first thread is different from the second thread.
[0017] In some embodiments, the front end is a web page.
[0018] According to a third aspect of the embodiments of the present disclosure, there is provided a semantic annotation device applied to a server. The device includes: a first acquisition module, configured to acquire a scene video sequence of a three-dimensional scene sent by a front end; a reconstruction module, configured to perform scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain description information of the reconstructed three-dimensional scene; and a first sending module, configured to send the description information of the three-dimensional scene to the front end to obtain a semantic annotation result of the three-dimensional scene returned by the front end, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene.
[0019] In some embodiments, each frame image in the scene video sequence includes an R-channel image, a G-channel image, a B-channel image, and a depth image of the three-dimensional scene.
[0020] In some embodiments, the first sending module includes: a reconstruction unit, configured to perform scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain a plurality of grids corresponding to the three-dimensional scene; and an acquisition unit, configured to respectively acquire description information of each grid in the plurality of grids, where the description information of the three-dimensional scene includes the description information of each grid.
[0021] In some embodiments, the device further includes: a second acquisition module, configured to acquire a semantic label of each point in the three-dimensional scene; a generation module, configured to generate a semantic label of each grid according to the semantic labels of each point in each grid in the plurality of grids after obtaining the plurality of grids corresponding to the three-dimensional scene; wherein each grid includes at least one point in the three-dimensional scene; and a third sending module, configured to send the semantic label of each grid to the front end to obtain a correction result of the semantic annotation result based on the semantic label of each grid at the front end.
[0022] In some embodiments, the apparatus further comprises: a third acquisition module, configured to acquire the semantic annotation result returned by the front end after sending the description information of the three-dimensional scene to the front end; and a projection module, configured to project the semantic annotation result onto each frame image of the scene video sequence.
[0023] In some embodiments, the number of the front ends is plural.
[0024] In some embodiments, the apparatus further comprises: a fourth acquisition module, configured to respectively acquire the semantic annotation results of each of the plural front ends; and a storage module, configured to store the semantic annotation results of each front end according to the scenes corresponding to the semantic annotation results of each front end.
[0025] According to a fourth aspect of the embodiments of the present disclosure, there is provided a semantic annotation apparatus, which is applied to a front end. The apparatus comprises: a second sending module, configured to send a scene video sequence of a three-dimensional scene to a server, so that the server performs scene reconstruction on the three-dimensional scene according to the scene video sequence; a receiving module, configured to receive the description information of the three-dimensional scene returned by the server; and a returning module, configured to generate the three-dimensional scene according to the description information of the three-dimensional scene and return the semantic annotation result of the three-dimensional scene to the server, wherein the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene.
[0026] In some embodiments, the description information of the reconstructed three-dimensional scene includes the description information of each of a plurality of meshes in the three-dimensional scene; the returning module comprises: an aggregation unit, configured to aggregate the plurality of meshes according to the description information of each of the plurality of meshes to obtain at least one aggregated mesh, and each of the at least one aggregated meshes corresponds to an object in the three-dimensional scene; and an annotation unit, configured to perform semantic annotation on each of the at least one aggregated meshes to obtain the semantic annotation result of the three-dimensional scene.
[0027] In some embodiments, the apparatus further comprises: a hiding module, configured to hide at least one first semantic annotation result in the semantic annotation result; and / or a display module, configured to display at least one first semantic annotation result that has been hidden.
[0028] In some embodiments, the hiding module comprises: a generating unit, configured to generate a set of cut surfaces for each of the at least one first semantic annotation results, wherein a set of cut surfaces of a first semantic annotation result wraps the first semantic annotation result; and a hiding unit, configured to hide each of the at least one first semantic annotation results through the set of cut surfaces of each of the at least one first semantic annotation results.
[0029] In some embodiments, the annotation unit includes: a receiving subunit, configured to receive a selection instruction for each of the at least one aggregated grid; a generating subunit, configured to generate a bounding box for each of the aggregated grids according to the selection instruction corresponding to each of the aggregated grids, where the bounding box of an aggregated grid includes each point in the aggregated grid, and the bounding box is a bounding box or a convex hull; and an annotating subunit, configured to perform semantic annotation on the points in the bounding box of each of the aggregated grids.
[0030] In some embodiments, generating the three-dimensional scene according to the description information of the three-dimensional scene is executed by a first thread; and obtaining the semantic annotation result of the three-dimensional scene is executed by a second thread; wherein, the first thread is different from the second thread.
[0031] In some embodiments, the front end is a web page.
[0032] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any of the embodiments is implemented.
[0033] According to a sixth aspect of the embodiments of the present disclosure, there is provided a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the program, the method described in any of the embodiments is implemented.
[0034] According to a seventh aspect of the embodiments of the present disclosure, there is provided a semantic annotation system, where the semantic annotation system includes: at least one front end; and a server; each front end in the at least one front end is configured to send a scene video sequence of a three-dimensional scene to the server, so that the server performs scene reconstruction on the three-dimensional scene according to the scene video sequence, receive the description information of the three-dimensional scene returned by the server after performing scene reconstruction, generate the three-dimensional scene according to the description information of the three-dimensional scene returned by the server, and return the semantic annotation result of the three-dimensional scene to the server, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene; and the server is configured to perform scene reconstruction on the three-dimensional scene corresponding to each front end according to the scene video sequence sent by each front end in the at least one front end, and obtain the description information of the three-dimensional scene corresponding to each front end.
[0035] In the embodiments of the present disclosure, the scene video sequence of the three-dimensional scene is sent to the server through the front end. The server performs scene reconstruction to obtain the description information of the reconstructed three-dimensional scene. Then, the front end generates the three-dimensional scene locally according to the description information of the reconstructed three-dimensional scene and returns the semantic annotation result of the three-dimensional scene to the server. Among them, the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene. By accurately reconstructing the three-dimensional scene and then performing semantic annotation based on the description information of the reconstructed three-dimensional scene, the annotation accuracy is improved.
[0036] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The drawings herein are incorporated into the specification and constitute a part of this specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0038] Figure 1 It is a flowchart of the semantic annotation method according to the embodiments of the present disclosure.
[0039] Figure 2 It is a schematic diagram of the processing flow of the server according to the embodiments of the present disclosure.
[0040] Figure 3 It is a flowchart of the semantic annotation method according to other embodiments of the present disclosure.
[0041] Figure 4 It is a schematic diagram of the annotation interface according to the embodiments of the present disclosure.
[0042] Figures 5A to 5C It is a schematic diagram of the annotation process according to the embodiments of the present disclosure.
[0043] Figure 6 It is a schematic diagram of the operation guide according to the embodiments of the present disclosure.
[0044] Figure 7 It is a block diagram of the semantic annotation measurement device according to the embodiments of the present disclosure.
[0045] Figure 8 It is a block diagram of the semantic annotation measurement device according to other embodiments of the present disclosure.
[0046] Figure 9 It is a schematic diagram of the computer device according to the embodiments of the present disclosure.
[0047] Figure 10 It is a schematic diagram of the semantic annotation system according to the embodiments of the present disclosure.
[0048] Figure 11 It is an interaction diagram of the semantic annotation system according to the embodiments of the present disclosure. Detailed implementation manners
[0049] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of apparatuses and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0050] The terms used in the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. The singular forms "a", "the", and "said" used in the present disclosure and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items. In addition, the term "at least one" herein means any one of a plurality or any combination of at least two of a plurality.
[0051] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0052] In order to enable those skilled in the art to better understand the technical solutions in the embodiments of the present disclosure and make the above-mentioned objects, features, and advantages of the embodiments of the present disclosure more obvious and understandable, the technical solutions in the embodiments of the present disclosure will be further described in detail below with reference to the drawings.
[0053] With the development of technology, artificial intelligence has been more and more widely applied. Computer vision is an important part of artificial intelligence, and semantic annotation of scenes is of great significance to computer vision. Semantic annotation is used to solve the problem of which target each point in the scene belongs to. At present, semantic annotation has extensive requirements and applications in many fields such as geographic information systems, unmanned vehicle driving, medical image analysis, and robots. For example, in the field of geographic information systems, roads, rivers, crops, buildings, etc. can be automatically identified through semantic annotation. In the field of unmanned driving, vehicles can automatically avoid pedestrians and obstacles through semantic annotation. In the field of intelligent medicine, the main applications of semantic annotation include tumor image annotation, dental caries diagnosis, etc.
[0054] However, most traditional semantic annotation methods are for two-dimensional images and lack support for semantic annotation of three-dimensional scenes. In actual production and life, there is a significant demand for datasets of three-dimensional semantic annotation.
[0055] Based on this, embodiments of the present disclosure provide a semantic annotation method, which is applied to a server. As Figure 1 shown, the method may include:
[0056] Step 101: Obtain a scene video sequence of a three-dimensional scene sent by the front end;
[0057] Step 102: Perform scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain description information of the reconstructed three-dimensional scene;
[0058] Step 103: Send the description information of the three-dimensional scene to the front end to obtain the semantic annotation result of the three-dimensional scene returned by the front end, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene.
[0059] In step 101, based on the scene video sequence, color information and position information of each point in the three-dimensional scene can be obtained. In some embodiments, each frame image in the scene video sequence includes an R-channel image, a G-channel image, a B-channel image, and a depth image of the three-dimensional scene. Among them, the R-channel image, the G-channel image, and the B-channel image are used to determine the color information of each point in the three-dimensional scene, and the depth image is used to determine the position information of each point in the three-dimensional scene.
[0060] In step 102, the color information and position information of each point in the three-dimensional scene can be determined according to the scene video sequence, so as to perform scene reconstruction on the three-dimensional scene to obtain description information of the reconstructed three-dimensional scene. Specifically, the three-dimensional scene can be reconstructed according to the scene video sequence to obtain multiple meshes corresponding to the three-dimensional scene; description information of each mesh in the multiple meshes is respectively obtained, where the description information of the three-dimensional scene includes the description information of each mesh. When performing scene reconstruction on the three-dimensional scene, multiple scene meshes of the entire three-dimensional scene can be obtained, and then, each scene mesh can be pre-annotated to obtain multiple meshes corresponding to the three-dimensional scene. Each scene mesh can be pre-annotated as one or more meshes. In this way, the annotation accuracy can be improved.
[0061] The description information of each grid may include the position information and color information of the points that make up each grid, and may also include at least one of the normal vector information and texture information of the points that make up each grid. The server may generate a json file according to the above description information. The json file may include multiple arrays, and each array is used to represent the description information of a grid. Different elements in the array respectively represent the position information, color information, normal vector information, and texture information. On the annotation plane (such as the ground, ceiling, etc.), by using the normal vector information, it helps to determine the orientation of the plane; by using the texture information, more details can be presented in the scene, making the scene more realistic and helping to perform semantic annotation on the scene. In practical applications, the description information may also include other information according to actual needs, which will not be elaborated here.
[0062] In step 103, the description information of the three-dimensional scene may be sent to the front end, and the front end automatically performs semantic annotation on the three-dimensional scene based on the description information of the three-dimensional scene to obtain the semantic annotation result of the three-dimensional scene. Alternatively, the user may perform semantic annotation on the three-dimensional scene based on the description information of the three-dimensional scene at the front end to obtain the semantic annotation result of the three-dimensional scene. If the description information of each grid is obtained in step 102, the description information of each grid may be sent to the front end in this step. The front end may generate the three-dimensional scene locally according to the description information.
[0063] In some embodiments, in order to make the annotation result more accurate, the server may also obtain the semantic label of each point in the three-dimensional scene. The semantic label of each point is used to predict the semantics of each point; after obtaining multiple grids corresponding to the three-dimensional scene, according to the semantic labels of each point in each grid among the multiple grids, generate the semantic label of each grid; send the semantic label of each grid to the front end to obtain the correction result of the semantic annotation result based on the semantic label of each grid at the front end. The correction result may be automatically obtained by the front end correcting the semantic annotation result, or may be obtained by the user performing a correction operation on the semantic annotation result at the front end.
[0064] Among them, the semantic label may be carried in the description information and sent to the front end together with the description information, or the description information may be sent first and then the semantic label. In some embodiments, the server may input the reconstructed three-dimensional scene into a pre-trained neural network to obtain the semantic label of each point in the three-dimensional scene output by the neural network. By outputting the semantic label of each point in the three-dimensional scene through the neural network, higher accuracy can be obtained.
[0065] After obtaining the semantic annotation result, the front end can return the semantic annotation result to the server. The server can save the semantic annotation result for the front end to continue semantic annotation on the unannotated 3D scene. The semantic annotation result may include the objects already annotated in the 3D scene and the objects not annotated in the 3D scene.
[0066] Furthermore, the front end can also export the result of the most recent annotation from the server. Specifically, the server can receive an export instruction sent by a first front end among multiple front ends; in response to the export instruction, send the saved semantic annotation result of the first front end to the first front end for display. After the first front end exports the result of the most recent annotation, it can continue to annotate the result of the most recent annotation.
[0067] In some embodiments, the server can also project the semantic annotation result of the 3D scene onto a two-dimensional plane. Specifically, after sending the description information of the 3D scene to the front end, the server can obtain the semantic annotation result returned by the front end; project the semantic annotation result onto each frame image of the scene video sequence, thereby obtaining the annotation result on the two-dimensional plane. When projecting, the homography matrix of the image acquisition device can be determined according to the internal and external parameters of the image acquisition device that acquires the scene video sequence, and then the semantic annotation result can be projected according to the homography matrix.
[0068] In practical applications, the above front end can be a web page. Most traditional annotation tools are client software, which requires users to configure relevant environments on the deployment nodes, and the usage cost is relatively high. The embodiment of the present disclosure adopts a web-based annotation tool, making the annotation tool more convenient and supporting multi-person online annotation, which can improve the annotation speed.
[0069] In some embodiments, the number of front ends can be multiple. Specifically, the server can respectively obtain the scene video sequences of the 3D scenes sent by multiple front ends, and perform scene reconstruction on the 3D scenes of each front end according to the scene video sequence sent by each front end, so as to obtain the description information of the reconstructed 3D scene of each front end. Then, send the description information corresponding to each front end to each front end to obtain the semantic annotation result of the 3D scene of each front end at each front end.
[0070] For example, assume the number of front-ends is 2. Then the server can obtain the scene video sequence 1 sent by front-end 1, perform scene reconstruction on the 3D scene 1 of front-end 1 based on the scene video sequence 1 to obtain the description information of the 3D scene 1, and send the description information of the 3D scene 1 to front-end 1 to generate the semantic annotation result of the 3D scene 1 at front-end 1, where the semantic annotation result of the 3D scene 1 is generated based on the description information of the 3D scene 1. Similarly, the server can also obtain the scene video sequence 2 sent by front-end 2, and then obtain the description information of the 3D scene 2 of front-end 2 in the same way to generate the semantic annotation result of the 3D scene 2 at front-end 2. The above example is only for illustrative purposes. In practical applications, the number of front-ends is not limited to this. The annotation method when the number of front-ends is greater than 2 is similar to the above example and will not be elaborated here.
[0071] In the case where the number of front-ends is multiple, different front-ends can perform semantic annotation on the same scene simultaneously, or on different scenes. For example, front-end 1 and front-end 2 can perform semantic annotation on the 3D scene 1 simultaneously. At the same time, front-end 3 can perform semantic annotation on the 3D scene 2. Each front-end can send its respective semantic annotation result to the server. The semantic annotation result returned by each front-end can carry the identification information of the front-end, which is used by the server to determine which front-end performs semantic annotation on each 3D scene. In the case where the number of front-ends is multiple, the server can return the saved annotation results to each front-end according to the identification information of each front-end, so that the front-end can export the annotation results and continue to perform semantic annotation on the scene.
[0072] In some embodiments, the server can also return the annotation result of one front-end to another front-end that performs semantic annotation on the same scene, so that the other front-end can refresh the annotation result. Specifically, the server can receive a refresh instruction sent by a first front-end among the multiple front-ends; in response to the refresh instruction, obtain the saved second semantic annotation results of at least one second front-end among the multiple front-ends; and send the second semantic annotation results of the at least one second front-end to the first front-end, so that the first front-end can refresh the first semantic result displayed locally according to the second semantic annotation results. Wherein, the first front-end and the second front-end annotate the same 3D scene. In some embodiments, the semantic annotation result returned by each front-end can carry the identification information of the 3D scene annotated by the front-end, which is used by the server to determine the scenes annotated by each front-end. When the server receives the refresh instruction from the first front-end, it can determine the second semantic annotation results of the second front-ends that annotate the same scene as the first front-end among the saved semantic annotation results according to the identification information of each 3D scene.
[0073] Such asFigure 2 As shown, it is a schematic diagram of the specific processing flow of the server in an embodiment of the present disclosure. In this embodiment, first, the server acquires an RGB-D video sequence, that is, a video sequence in which each frame image includes an R-channel image, a G-channel image, a B-channel image, and a depth image. Scene reconstruction is performed based on the RGB-D video sequence, and then the reconstructed scene is pre-annotated to obtain multiple meshes and their description information. This description information can be output to a database (such as a MongoDB database) for storage, for the front end to export and generate semantic annotation results at the front end.
[0074] As Figure 3 shown, an embodiment of the present disclosure also provides a semantic annotation method, which is applied to the front end. The method may include:
[0075] Step 301: Send the scene video sequence of the three-dimensional scene to the server, so that the server performs scene reconstruction on the three-dimensional scene according to the scene video sequence;
[0076] Step 302: Receive the description information of the three-dimensional scene returned by the server after scene reconstruction;
[0077] Step 303: Generate the three-dimensional scene according to the description information of the three-dimensional scene, and return the semantic annotation result of the three-dimensional scene to the server, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene.
[0078] In some embodiments, the description information of the three-dimensional scene includes the description information of each mesh in the multiple meshes in the three-dimensional scene. For example, it includes the position information and color information of the points constituting each mesh; for another example, it may also include the normal vector information and texture information of the points constituting each mesh. The front end can aggregate the multiple meshes according to the description information of each mesh in the multiple meshes to obtain at least one aggregated mesh, and each aggregated mesh in the at least one aggregated mesh corresponds to an object in the three-dimensional scene; perform semantic annotation on each aggregated mesh in the at least one aggregated mesh to obtain the semantic annotation result of the three-dimensional scene. For example, if the objects corresponding to meshes 1 to 5 are all tables and the objects corresponding to meshes 6 to 10 are all chairs, then meshes 1 to 5 are merged into one aggregated mesh, and meshes 6 to 10 are merged into another aggregated mesh.
[0079] By coloring the aggregated meshes, the aggregated meshes of the same color form the semantic annotation of an object (such as a person, an animal, or an item). During the semantic annotation process, a three-dimensional box parallel to the three-dimensional coordinate axes of each object can be generated, and the three-dimensional box corresponding to the currently annotated aggregated mesh is used to quickly color the selected aggregated mesh.
[0080] In some embodiments, the method further includes: hiding at least one first semantic annotation result in the semantic annotation result. By hiding the annotated semantic annotation results, the occlusion of the unannotated meshes by the annotated semantic annotation results can be reduced, which helps to confirm the edge details of the unannotated meshes. Specifically, a set of cutting planes can be generated for each first semantic annotation result in the at least one first semantic annotation result; and each first semantic annotation result is hidden through the set of cutting planes of each first semantic annotation result. The cutting plane has no color and is not displayed in the scene. Its function is to occlude all objects behind this plane and only display the objects in front of the cutting plane. Among them, a set of cutting planes of a first semantic annotation result wraps the first semantic annotation result. For example, a cutting plane can be generated in front of, behind, to the left, to the right, above, and below a first semantic annotation result. The above six cutting planes form a cuboid that wraps the first semantic annotation result. The objects outside the cuboid are hidden and only the inside is visible. Using cutting planes to hide the content that you don't want to see currently can reduce some occlusion situations.
[0081] In some embodiments, the method further includes: displaying at least one first semantic annotation result that has been hidden. By displaying at least one first semantic annotation result that has been hidden, it helps to observe and modify the overall situation of the annotation.
[0082] During annotation, the front end can also automatically perform semantic annotation on the meshes within the convex hull or bounding box by calculating the convex hull or bounding box of the aggregated meshes to speed up the annotation speed. Specifically, the front end can receive a selection instruction for each aggregated mesh in the at least one aggregated mesh; generate a bounding box for each aggregated mesh according to the selection instruction corresponding to each aggregated mesh; and perform semantic annotation on the points in the bounding box of each aggregated mesh. Among them, the bounding box of an aggregated mesh includes each point in the aggregated mesh, and the bounding box is a bounding box or a convex hull.
[0083] A bounding box is used to approximately replace a complex geometric object, which is a geometric body slightly larger in volume and simpler in characteristics than the geometric object. The convex hull is a convex polygon used to approximate the contour of a geometric object. In different situations, the bounding box or the convex hull can be flexibly selected for semantic annotation. The convex hull is closer than the bounding box, and the enclosed space formed is closer to the shape of the enclosed object. Generally speaking, when the object to be annotated is adjacent to other objects relatively closely, in order to avoid misannotating other objects, the convex hull can be used; while when the object to be annotated is relatively independent, the bounding box can be used. In practical applications, in addition to selecting the bounding box or the convex hull according to the distance between the objects to be annotated, other conditions can also be used for selection, which will not be elaborated here.
[0084] In some embodiments, the operation history during the annotation process can also be recorded, enabling the user to promptly undo or redo operations when the operations are inappropriate. Specifically, the operation results of each operation during the annotation process can be cached, and when a user sends an undo instruction, the operation result cached last time is called and displayed.
[0085] The basic interface of the front end of the embodiments of the present disclosure is as Figure 4 shown, including a category bar, an operation sub-interface, a menu bar, and an option bar. The category bar is used to manage the category information of the annotation. For example, the categories of the annotation objects can be added or deleted, and the labels of the annotation objects can be set, etc. The operation sub-interface includes a three-dimensional scene model, and the three-dimensional scene model includes a plurality of meshes. Operations can be performed on this interface to generate an aggregated mesh and perform semantic annotation on the aggregated mesh. The menu bar includes a "Save" option for saving the annotation result, a "Submit" option for exporting the annotation result, a "Help" option for viewing the operation guide, and options for viewing the previous task or the next task, etc. The option bar is used to set the parameters used during the annotation process. For example, the background color, the shape of the bounding box, etc. The above interface is only for illustrative purposes and is not used to limit the present disclosure. In addition to the above-mentioned various parts, according to actual needs, other parts can also be included on the interface of the front end. For example, switching the display language to Chinese or English, etc., will not be elaborated here.
[0086] When performing semantic annotation, the user can select the category of the annotation in the above-mentioned category bar, and then select the aggregated mesh on the operation sub-interface, and then semantic annotation can be performed on the selected aggregated mesh. The selection method can be an input method of mouse click, or a combined input method of mouse click and shortcut keys. For example, the aggregated mesh on the operation sub-interface can be held down Ctrl and clicked to perform semantic annotation on it. Further, the annotation of the annotated aggregated mesh can also be cancelled. For example, the annotated aggregated mesh can be held down Shift+Ctrl and clicked to cancel its annotation.
[0087] As Figure 5A shown, after the annotation object is selected, the front end can automatically generate the bounding box or convex hull of the object. Further, the shape of the bounding box can also be selected to obtain different bounding sizes, thereby changing the annotation range. As Figure 5B shown, the annotated object is filled with a certain color, which can be pre-specified or randomly selected. As Figure 5C shown, when performing semantic annotation on other objects, the annotated objects can be hidden, and only the objects to be annotated are displayed, thereby avoiding the occlusion of the unannotated objects by the annotated objects and making the operation interface look more concise.
[0088] To improve the annotation efficiency, shortcut keys can be preset to correspond to various operations during the annotation process. For example, Figure 6 as shown, the operation guide can be invoked to obtain the shortcut keys corresponding to various operations. For example, pressing the G key can automatically perform semantic annotation on the grids within the convex hull or bounding box, pressing the M key can transform the annotation scene, pressing the C key can hide or display the selected annotated objects, pressing the X key can hide or display all annotated objects, and pressing the ENTER key can cancel the selection, etc.
[0089] In some embodiments, generating the three-dimensional scene locally according to the description information of the three-dimensional scene is executed by a first thread; the semantic annotation result of the three-dimensional scene is obtained by a second thread; wherein, the first thread is different from the second thread. Using different threads for scene loading and obtaining semantic annotation results can prevent the page from freezing or even crashing when loading a large scene, making the loading process smoother.
[0090] The front end of the embodiments of the present disclosure can be a web page. Most traditional annotation tools are client software. The embodiments of the present disclosure adopt a web-based annotation tool, which brings great cross-platform convenience, can be used without installation, and supports multi-person online annotation.
[0091] The embodiments of the present disclosure have the following advantages:
[0092] (1) It solves the defect that traditional semantic annotation tools can only perform semantic annotation on two-dimensional images, realizes semantic annotation of three-dimensional scenes, and has relatively high annotation accuracy.
[0093] (2) Developed based on the web, it does not depend on the platform and installation environment, has good convenience, and eliminates the installation cost. And this tool can support multi-person online annotation and save the annotation progress before exiting.
[0094] (3) It has good expandability and can easily add corresponding functions as needed.
[0095] (4) It can calculate the convex hull or bounding box of all objects within a label according to the user's selection, and then automatically label all the objects within the convex hull or bounding box with the same label, improving the usability and annotation speed.
[0096] (5) It can check whether the annotation is complete by hiding the annotated objects, solving the problem of line-of-sight occlusion when the scene is complex.
[0097] (6) Using multi-threaded progressive loading of the scene can prevent the page from freezing or even crashing when loading a large scene, making the loading process smoother and improving the user experience.
[0098] (7) Multilingual support is added, along with a relatively detailed operation guide, making the operation simpler, and users who don't understand the principle can also use it conveniently.
[0099] (8) It can record the operation history, enabling users to easily roll back or redo previous operations when the operations are improper.
[0100] (9) The code structure of this annotation tool is clear and has strong scalability, facilitating the addition of new functions.
[0101] (9) Those skilled in the art can understand that in the above method of the specific implementation manner, the writing order of each step does not mean a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined according to its function and possible internal logic.
[0102] As Figure 7 shown, the present disclosure also provides a device, and the device includes:
[0103] A first acquisition module 701, configured to acquire a scene video sequence of a three-dimensional scene sent by the front end;
[0104] A reconstruction module 702, configured to perform scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain description information of the reconstructed three-dimensional scene;
[0105] A first sending module 703, configured to send the description information of the three-dimensional scene to the front end to obtain a semantic annotation result of the three-dimensional scene returned by the front end, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene.
[0106] In some embodiments, each frame image in the scene video sequence includes an R-channel image, a G-channel image, a B-channel image, and a depth image of the three-dimensional scene.
[0107] In some embodiments, the first sending module includes: a reconstruction unit, configured to perform scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain a plurality of meshes corresponding to the three-dimensional scene; an acquisition unit, configured to respectively acquire description information of each mesh in the plurality of meshes, where the description information of the three-dimensional scene includes the description information of each mesh.
[0108] In some embodiments, the apparatus further comprises: a second acquisition module, configured to acquire the semantic label of each point in the three-dimensional scene; a generation module, configured to generate the semantic label of each grid according to the semantic labels of the points in each grid among the multiple grids after obtaining the multiple grids corresponding to the three-dimensional scene, where each grid includes at least one point in the three-dimensional scene; and a third transmission module, configured to transmit the semantic label of each grid to the front end, so as to obtain, at the front end, a correction result of the semantic annotation result based on the semantic label of each grid.
[0109] In some embodiments, the apparatus further comprises: a third acquisition module, configured to acquire the semantic annotation result returned by the front end after transmitting the description information of the three-dimensional scene to the front end; and a projection module, configured to project the semantic annotation result onto each frame image of the scene video sequence.
[0110] In some embodiments, the number of the front ends is multiple.
[0111] In some embodiments, the apparatus further comprises: a fourth acquisition module, configured to respectively acquire the semantic annotation results of each of the multiple front ends; and a storage module, configured to store the semantic annotation results of each front end according to the scene corresponding to the semantic annotation result of each front end.
[0112] As Figure 8 shown, the present disclosure further provides an apparatus, where the apparatus comprises:
[0113] A second transmission module 801, configured to transmit a scene video sequence of a three-dimensional scene to a server, so that the server performs scene reconstruction on the three-dimensional scene according to the scene video sequence;
[0114] A reception module 802, configured to receive the description information of the three-dimensional scene returned by the server;
[0115] A return module 803, configured to generate the three-dimensional scene according to the description information of the three-dimensional scene and return the semantic annotation result of the three-dimensional scene to the server, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene.
[0116] In some embodiments, the description information of the reconstructed three-dimensional scene includes the description information of each grid in the multiple grids in the three-dimensional scene; the returning module includes: an aggregating unit, configured to aggregate the multiple grids according to the description information of each grid in the multiple grids to obtain at least one aggregated grid, and each aggregated grid in the at least one aggregated grid corresponds to an object in the three-dimensional scene; an annotating unit, configured to perform semantic annotation on each aggregated grid in the at least one aggregated grid to obtain the semantic annotation result of the three-dimensional scene.
[0117] In some embodiments, the apparatus further includes: a hiding module, configured to hide at least one first semantic annotation result in the semantic annotation result; and / or a displaying module, configured to display at least one first semantic annotation result that has been hidden.
[0118] In some embodiments, the hiding module includes: a generating unit, configured to generate a set of sections for each first semantic annotation result in the at least one first semantic annotation result, wherein a set of sections of a first semantic annotation result wraps the first semantic annotation result; a hiding unit, configured to hide each first semantic annotation result through the set of sections of each first semantic annotation result.
[0119] In some embodiments, the annotating unit includes: a receiving subunit, configured to receive a selection instruction for each aggregated grid in the at least one aggregated grid; a generating subunit, configured to generate a bounding box for each aggregated grid according to the selection instruction corresponding to each aggregated grid, and the bounding box of an aggregated grid includes each point in the aggregated grid, and the bounding box is a bounding box or a convex hull; an annotating subunit, configured to perform semantic annotation on the points in the bounding box of each aggregated grid.
[0120] In some embodiments, generating the three-dimensional scene locally according to the description information of the three-dimensional scene is executed by a first thread; the semantic annotation result of the three-dimensional scene is obtained by a second thread; wherein, the first thread is different from the second thread.
[0121] In some embodiments, the front end is a web page.
[0122] In some embodiments, the functions or modules included in the apparatus provided in the embodiments of the present disclosure can be used to execute the methods described in the method embodiments above, and the specific implementation can refer to the description of the method embodiments above. For the sake of brevity, it will not be repeated here.
[0123] The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the objectives of the solutions in this specification. A person of ordinary skill in the art can understand and implement them without creative efforts.
[0124] Correspondingly, an embodiment of the present disclosure also provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the method described in any one of the embodiments.
[0125] Figure 9 FIG. shows a more specific schematic diagram of the hardware structure of a computer device provided by an embodiment of this specification. The device may include: a processor 901, a memory 902, an input / output interface 903, a communication interface 904, and a bus 905. Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.
[0126] The processor 901 can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this specification.
[0127] The memory 902 can be implemented in forms such as a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, etc. The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and called and executed by the processor 901.
[0128] The input / output interface 903 is used to connect to an input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device can include a display, a speaker, a vibrator, an indicator light, etc.
[0129] The communication interface 904 is used to connect to a communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. The communication module can communicate through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0130] The bus 905 includes a path for transmitting information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904).
[0131] It should be noted that although the above device only shows the processor 901, the memory 902, the input / output interface 903, the communication interface 904, and the bus 905, in the specific implementation process, the device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0132] The embodiments of the present disclosure also provide a semantic annotation system, as Figure 10 shown, the semantic annotation system may include:
[0133] At least one front end 1001; and a server 1002;
[0134] Each front end in the at least one front end 1001 is used to send the scene video sequence of the three-dimensional scene to the server 1002, so that the server 1002 performs scene reconstruction on the three-dimensional scene according to the scene video sequence, receives the description information of the three-dimensional scene returned after the server 1002 performs scene reconstruction, generates the three-dimensional scene according to the description information of the three-dimensional scene returned by the server 1002, and returns the semantic annotation result of the three-dimensional scene to the server 1002, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene;
[0135] The server 1002 is used to perform scene reconstruction on the three-dimensional scene corresponding to each front end according to the scene video sequence sent by each front end in the at least one front end 1001, and obtain the description information of the reconstructed three-dimensional scene corresponding to each front end.
[0136] The figure shows the case where the number of front ends is n, where n is a positive integer. For example, n can be 1, or 2 or an integer greater than 2. Each front end in the embodiments of the present disclosure can be a web page. Through web-based design, it can meet the simultaneous online operation of multiple users and improve the annotation efficiency.
[0137] Since the interaction mode between each front end 1001 and the server 1002 is similar, see Figure 11 , here only one of the front ends is taken as an example to illustrate the interaction mode between the two.
[0138] In step 1101, the front end 1001 sends a video sequence to the server 1002.
[0139] In step 1102, the server 1002 performs scene reconstruction based on the video sequence to obtain the description information of multiple grids in the scene.
[0140] In step 1103, the server 1002 sends the description information to the front end 1001.
[0141] In step 1104, the front end 1001 performs grid aggregation on the multiple grids according to the description information to obtain the aggregated grid corresponding to each object.
[0142] In step 1105, the front end 1001 automatically performs semantic annotation on each aggregated grid, or the user performs semantic annotation on each aggregated grid at the front end, and each aggregated grid can be annotated with different colors.
[0143] In step 1106, the front end 1001 returns the semantic annotation result to the server 1002.
[0144] In step 1107, the server 1002 saves the semantic annotation result returned by the front end 1001.
[0145] For the specific embodiments of the front end 1001 in the embodiments of the present disclosure, see the foregoing method embodiments applied to the front end. For the specific embodiments of the server 1002 in the embodiments of the present disclosure, see the foregoing method embodiments applied to the server, which will not be elaborated here.
[0146] The embodiments of this specification also provide a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the method described in any of the foregoing embodiments is implemented.
[0147] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information accessible by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0148] From the description of the above embodiments, those skilled in the art can clearly understand that the embodiments of this specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions of the embodiments of this specification, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of this specification.
[0149] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer, and the specific form of the computer can be a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email transceiver, game console, tablet computer, wearable device, or a combination of any several of these devices.
[0150] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, they are described relatively simply. For the relevant parts, reference can be made to the descriptions in the method embodiments. The apparatus embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separated. When implementing the solutions of the embodiments of this specification, the functions of the modules can be realized in one or more software and / or hardware. It is also possible to select some or all of the modules according to actual needs to achieve the purpose of the solutions of this embodiment. A person of ordinary skill in the art can understand and implement it without creative efforts.
[0151] The above are only the specific implementation manners of the embodiments of this specification. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principles of the embodiments of this specification, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the embodiments of this specification.
Claims
1. A semantic annotation method, characterized in that, Applied to a server, the method includes: Obtain a scene video sequence of a three-dimensional scene sent by the front end; Perform scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain description information of the reconstructed three-dimensional scene; the description information is used for the front end to generate the three-dimensional scene locally; Send the description information of the three-dimensional scene to the front end to obtain the semantic annotation result of the three-dimensional scene returned by the front end, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene; The description information of the three-dimensional scene includes the description information of each grid among multiple grids in the three-dimensional scene; the semantic annotation result of the three-dimensional scene is obtained based on the following method: Aggregate the multiple grids according to the description information of each grid among the multiple grids to obtain at least one aggregated grid, and each aggregated grid in the at least one aggregated grid corresponds to an object in the three-dimensional scene; Perform semantic annotation on each aggregated grid in the at least one aggregated grid to obtain the semantic annotation result of the three-dimensional scene.
2. The method according to claim 1, characterized in that, Each frame image in the scene video sequence includes an R-channel image, a G-channel image, a B-channel image, and a depth image of the three-dimensional scene.
3. The method according to claim 1, characterized in that, The performing scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain description information of the reconstructed three-dimensional scene includes: Perform scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain multiple grids corresponding to the three-dimensional scene; Respectively obtain the description information of each grid among the multiple grids, where the description information of the three-dimensional scene includes the description information of each grid.
4. The method according to claim 3, characterized in that, The method further includes: Obtain the semantic label of each point in the three-dimensional scene; After obtaining the multiple grids corresponding to the three-dimensional scene, generate the semantic label of each grid according to the semantic labels of each point in each grid among the multiple grids; where each grid includes at least one point in the three-dimensional scene; Send the semantic label of each grid to the front end to obtain a correction result of the semantic annotation result obtained based on the semantic label of each grid at the front end.
5. The method according to claim 1, characterized in that The method further includes: After sending the description information of the three-dimensional scene to the front end, obtain the semantic annotation result returned by the front end; Project the semantic annotation result onto each frame image of the scene video sequence.
6. The method according to any one of claims 1 to 5, characterized in that The number of the front ends is multiple; the method further includes: Respectively obtain the semantic annotation results of each of the multiple front ends; Save the semantic annotation results of each front end according to the scenes corresponding to the semantic annotation results of each front end.
7. A semantic annotation method, characterized in that, Applied to the front end, the method includes: Send the scene video sequence of the three-dimensional scene to the server so that the server performs scene reconstruction on the three-dimensional scene according to the scene video sequence; Receive the description information of the three-dimensional scene returned after the server performs scene reconstruction; Generate the 3D scene locally according to the description information of the 3D scene, and return the semantic annotation result of the 3D scene to the server, where the semantic annotation result of the 3D scene is obtained based on the description information of the 3D scene; The description information of the 3D scene includes the description information of each grid in multiple grids in the 3D scene; the semantic annotation result of the 3D scene is obtained based on the following method: Aggregate the multiple grids according to the description information of each grid in the multiple grids to obtain at least one aggregated grid, and each aggregated grid in the at least one aggregated grid corresponds to an object in the 3D scene; Perform semantic annotation on each aggregated grid in the at least one aggregated grid to obtain the semantic annotation result of the 3D scene.
8. The method according to claim 7, wherein The method further includes: Hide at least one first semantic annotation result in the semantic annotation result; and / or Display at least one first semantic annotation result that has been hidden.
9. The method according to claim 8, characterized in that, The hiding of at least one first semantic annotation result in the semantic annotation result includes: Generate a set of sections for each first semantic annotation result in the at least one first semantic annotation result, where a set of sections of a first semantic annotation result wraps the first semantic annotation result; Hide each first semantic annotation result through a set of sections of each first semantic annotation result.
10. The method according to claim 7, wherein The performing of semantic annotation on each aggregated grid in the at least one aggregated grid includes: Receive a selection instruction for each aggregated grid in the at least one aggregated grid; Generate a bounding box for each aggregated grid according to the selection instruction corresponding to each aggregated grid, and the bounding box of an aggregated grid includes each point in the aggregated grid, and the bounding box is a bounding box or a convex hull; Perform semantic annotation on the points in the bounding box of each aggregated grid.
11. The method according to claim 7, wherein The generation of the 3D scene according to the description information of the 3D scene is executed by a first thread; The semantic annotation result of the 3D scene is obtained by a second thread; wherein, the first thread is different from the second thread.
12. The method according to any one of claims 7 to 11, characterized in that, The front end is a web page.
13. A semantic annotation device, characterized in that, Applied to a server, the device includes: A first acquisition module, configured to acquire a scene video sequence of a 3D scene sent by the front end; A reconstruction module, configured to perform scene reconstruction on the 3D scene according to the scene video sequence to obtain the description information of the reconstructed 3D scene; the description information is used for the front end to generate the 3D scene locally; A first sending module, configured to send the description information of the 3D scene to the front end to obtain the semantic annotation result of the 3D scene returned by the front end, where the semantic annotation result of the 3D scene is obtained based on the description information of the 3D scene; The description information of the 3D scene includes the description information of each grid in multiple grids in the 3D scene; the semantic annotation result of the 3D scene is obtained by the front end based on the following modules: An aggregation module, configured to aggregate the multiple meshes according to the description information of each mesh in the multiple meshes, so as to obtain at least one aggregated mesh, and each aggregated mesh in the at least one aggregated mesh corresponds to an object in the three-dimensional scene; A semantic annotation module, configured to perform semantic annotation on each aggregated mesh in the at least one aggregated mesh to obtain a semantic annotation result of the three-dimensional scene.
14. A semantic annotation device, characterized in that Applied to the front end, the device includes: A second sending module, configured to send a scene video sequence of a three-dimensional scene to a server, so that the server performs scene reconstruction on the three-dimensional scene according to the scene video sequence; A receiving module, configured to receive the description information of the three-dimensional scene returned by the server after performing scene reconstruction; A returning module, configured to generate the three-dimensional scene locally according to the description information of the three-dimensional scene, and return the semantic annotation result of the three-dimensional scene to the server, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene; The description information of the three-dimensional scene includes the description information of each mesh in the multiple meshes in the three-dimensional scene; the semantic annotation result of the three-dimensional scene is obtained based on the following modules: An aggregation module, configured to aggregate the multiple meshes according to the description information of each mesh in the multiple meshes, so as to obtain at least one aggregated mesh, and each aggregated mesh in the at least one aggregated mesh corresponds to an object in the three-dimensional scene; An annotation module, configured to perform semantic annotation on each aggregated mesh in the at least one aggregated mesh to obtain a semantic annotation result of the three-dimensional scene.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the method according to any one of claims 1 to 12.
16. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method according to any one of claims 1 to 12.
17. A semantic annotation system, characterized in that, The semantic annotation system includes: At least one front end; and a server; Each front end in the at least one front end is configured to send a scene video sequence of a three-dimensional scene to the server, so that the server performs scene reconstruction on the three-dimensional scene according to the scene video sequence to obtain the description information of the reconstructed three-dimensional scene; the description information is used for the front end to generate the three-dimensional scene locally; receive the description information of the three-dimensional scene returned by the server after performing scene reconstruction, generate the three-dimensional scene according to the description information of the three-dimensional scene returned by the server, and return the semantic annotation result of the three-dimensional scene to the server, where the semantic annotation result of the three-dimensional scene is obtained based on the description information of the three-dimensional scene; The server is configured to perform scene reconstruction on the three-dimensional scene corresponding to each front end according to the scene video sequence sent by each front end in the at least one front end, and obtain the description information of the three-dimensional scene corresponding to each front end; The description information of the three-dimensional scene includes the description information of each grid among a plurality of grids in the three-dimensional scene; the semantic annotation result of the three-dimensional scene is obtained based on the following method: according to the description information of each grid among the plurality of grids, the plurality of grids are aggregated to obtain at least one aggregated grid, and each aggregated grid in the at least one aggregated grid corresponds to an object in the three-dimensional scene; semantic annotation is performed on each aggregated grid in the at least one aggregated grid to obtain the semantic annotation result of the three-dimensional scene.
Citation Information
Patent Citations
Web-based semantic annotation system and Web-based semantic annotation method for high resolution SAR (synthetic aperture radar) image interpretation
CN102708167A
A method and system for interactive image text annotation
CN109299296A
Three-dimensional point cloud labeling method, device, equipment and storage medium
CN111062255A
Scene roaming method, system and device based on three-dimensional modeling and storage medium
CN111080799A
Construction method and device of three-dimensional semantic map, electronic equipment and storage medium
CN111190981A