A method, device, mobile terminal and storage medium for sample data management

By dividing sample data into different types and storing them in tables, encodings and files, and combining with corresponding indexing strategies, the problem of low sample data retrieval efficiency in the existing technology is solved, and efficient data management and retrieval is achieved.

CN114385863BActive Publication Date: 2025-05-30广州海格星航信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111545280.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-05-30
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

In the prior art, the sample data management method has the problem of low data retrieval efficiency, especially when processing a large number of visual sample data, the database query speed is slow and it is impossible to quickly search samples based on space-time information.

Method used

The sample data to be managed is divided into three categories: the first sample (sample category data and sample table information data) is stored in a table, the second sample (sample position data) is stored in an encoding, the third sample (picture data and labeled result data) is stored in a file, and the corresponding indexing strategy is configured to improve retrieval efficiency.

Benefits of technology

It realizes flexible and efficient storage and efficient retrieval of sample data, improves data retrieval efficiency, reduces database storage space, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114385863B_ABST
    Figure CN114385863B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, an apparatus, a mobile terminal and a storage medium for sample data management. The method includes: after obtaining the sample data to be managed, dividing the sample data to be managed into a first sample, a second sample and a third sample according to the data nature; storing the first sample in a tabular manner, storing the second sample in an encoded manner, and storing the third sample in a file manner. By adopting the embodiment of the present invention, the retrieval efficiency of sample data can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular, to a method and device for managing sample data, a mobile terminal, and a storage medium. Background Art

[0002] In recent years, with the progress of technology and the development of society, the application of artificial intelligence has advanced by leaps and bounds in many fields. The automation and intelligence technologies of vehicles such as driverless cars have been continuously improved, and a large number of samples of road target states under different visual conditions are required for learning and testing during the improvement of technology. At present, the management of large-scale road detection target samples is basically carried out by self-constructing a system, placing basic information in a database, and storing sample data in the form of a database or a single-node file library.

[0003] Storing sample data or annotation results in a database after encoding them into binary data or Base64 will greatly occupy the database space, resulting in an extremely slow speed of querying data in the database. Moreover, there is no effective storage method for location information, and it is impossible to quickly retrieve samples according to spatio-temporal information. In addition, storing samples or annotation results in a single node cannot better manage each sample server. When the disk occupancy is too large, regular maintenance is required, increasing the maintenance cost.

[0004] In summary, the method for managing sample data in the prior art has the problem of low data retrieval efficiency. Summary of the Invention

[0005] Embodiments of the present invention provide a method and device for managing sample data, a mobile terminal, and a storage medium, which improve the data retrieval efficiency.

[0006] The first aspect of the embodiments of the present application provides a method for managing sample data, including:

[0007] After obtaining the sample data to be managed, dividing the sample data to be managed into a first sample, a second sample, and a third sample according to the data nature;

[0008] Storing the first sample in a tabular manner, storing the second sample in an encoded manner, and storing the third sample in a file manner.

[0009] In a possible implementation manner of the first aspect, it further includes:

[0010] Configuring a table index strategy for the first sample, a grid index strategy for the second sample, and a directory file index strategy for the third sample.

[0011] In a possible implementation of the first aspect, the sample data to be managed is divided into a first sample, a second sample, and a third sample according to the data nature, specifically:

[0012] After obtaining the sample category data and the sample table information data in the data to be managed according to the data nature, the sample category data and the sample table information data are used as the first sample;

[0013] After obtaining the sample position data in the data to be managed according to the data nature, the sample position data is used as the second sample;

[0014] After obtaining the picture data and the annotation result data in the data to be managed according to the data nature, the picture data and the annotation result data are used as the third sample.

[0015] In a possible implementation of the first aspect, the second sample is stored in an encoded manner, specifically:

[0016] The sample position data is input into a preset grid model so that the preset grid model outputs three-dimensional space grid data according to the sample position data;

[0017] After the three-dimensional space grid data is converted into binary, a space grid code is generated and stored.

[0018] A second aspect of the embodiments of the present application provides a sample data management device, including: an acquisition module and a management module;

[0019] Among them, the acquisition module is used to obtain the sample data to be managed and then divide the sample data to be managed into a first sample, a second sample, and a third sample according to the data nature;

[0020] The management module is used to store the first sample in a table manner, store the second sample in an encoded manner, and store the third sample in a file manner.

[0021] In a possible implementation of the second aspect, it further includes:

[0022] Configure a table index strategy for the first sample, configure a grid index strategy for the second sample, and configure a directory file index strategy for the third sample.

[0023] In a possible implementation of the second aspect, the sample data to be managed is divided into a first sample, a second sample, and a third sample according to the data nature, specifically:

[0024] After obtaining the sample category data and the sample table information data in the data to be managed according to the data nature, the sample category data and the sample table information data are used as the first sample;

[0025] After obtaining the sample location data in the data to be managed according to the data nature, use the sample location data as the second sample;

[0026] After obtaining the picture data and the annotation result data in the data to be managed according to the data nature, use the picture data and the annotation result data as the third sample.

[0027] In a possible implementation manner of the second aspect, store the second sample in an encoded manner, specifically:

[0028] Input the sample location data into a preset grid model so that the preset grid model outputs three-dimensional space grid data according to the sample location data;

[0029] After performing binary conversion on the three-dimensional space grid data, generate a space grid code and store it.

[0030] A third aspect of the embodiments of the present application provides a mobile terminal, including a processor and a memory. The memory stores computer-readable program code, and when the processor executes the computer-readable program code, the steps of the above-mentioned sample data management method are implemented.

[0031] A fourth aspect of the embodiments of the present application provides a storage medium that stores computer-readable program code, and when the computer-readable program code is executed, the steps of the above-mentioned sample data management method are implemented.

[0032] Compared with the prior art, a sample data management method, device, mobile terminal and storage medium provided by the embodiments of the present invention, the method includes: after obtaining the sample data to be managed, divide the sample data to be managed into a first sample, a second sample and a third sample according to the data nature; store the first sample in a table manner, store the second sample in an encoded manner, and store the third sample in a file manner.

[0033] The beneficial effect thereof is that: in the embodiments of the present invention, different sample data are stored in different ways according to the data nature, which can flexibly realize the effective storage of sample data, and perform corresponding retrieval on different sample data according to different storage methods, which can effectively improve the retrieval efficiency of sample data.

[0034] Further, the second sample in the embodiments of the present invention is sample location data, and the third sample is picture data and annotation result data. Storing the sample location data in an encoded manner, that is, performing grid encoding and encoding compression on the sample location data, can reduce the storage of the sample location data. Storing the picture data and the annotation result data in a file manner, and storing the picture data and the annotation result data in a file server in a directory + file manner, can reduce the occupation of the database storage space.

[0035] Furthermore, the embodiments of the present invention configure different indexing strategies for sample data of different natures, which can further improve the retrieval efficiency. For example, for the retrieval of the second sample, since the second sample is stored in an encoded manner, the grid space data retrieval method (i.e., the grid indexing strategy) is used for retrieval, and the retrieval result can be obtained quickly.

[0036] Finally, the embodiments of the present invention adopt automatic configuration management for sample data, and specifically adopt a single-node or distributed method for sample data management. When the data volume of a single node has reached 20TB, the new data is automatically stored in the next node, and there is no need for manual file migration and maintenance, solving the problems of regular maintenance of servers and high maintenance costs in the prior art. Description of the Drawings

[0037] Figure 1 is a schematic flowchart of a sample data management method provided by an embodiment of the present invention;

[0038] Figure 2 is a schematic structural diagram of a sample data management device provided by an embodiment of the present invention. Detailed Embodiments

[0039] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0040] Refer to Figure 1 , which is a schematic flowchart of a sample data management method provided by an embodiment of the present invention, including S101 - S102:

[0041] S101: After obtaining the sample data to be managed, divide the sample data to be managed into a first sample, a second sample, and a third sample according to the data nature.

[0042] Preferably, the sample data to be managed is road target detection sample data, including: the location where the sample belongs, the real-time visual road condition information of the sample, and other sample-related data. Further, after obtaining the sample data to be managed, analyze the sample data to be managed, and relevant information such as annotation result data can be obtained.

[0043] In this embodiment, the step of dividing the sample data to be managed into a first sample, a second sample, and a third sample according to the data nature is specifically:

[0044] After obtaining the sample category data and the sample table information data in the data to be managed according to the data nature, use the sample category data and the sample table information data as the first sample;

[0045] After obtaining the sample position data in the data to be managed according to the data nature, use the sample position data as the second sample;

[0046] After obtaining the picture data and the annotation result data in the data to be managed according to the data nature, use the picture data and the annotation result data as the third sample.

[0047] S102: Store the first sample in the form of a table, store the second sample in the form of encoding, and store the third sample in the form of a file.

[0048] In this embodiment, the storing the first sample in the form of a table is specifically as follows:

[0049] After obtaining the sample category data in the data to be managed according to the data nature, save the sample category data as the first sample; if the sample category data does not exist, add a new sample category or obtain the sample category ID. The sample category data is stored in the form of a table, that is, the sample category data is stored in the form of a database table. Among them, the sample category data includes: the identifier of the sample category, the name of the sample category, the corresponding visual road conditions information (such as weather, lighting, paving, etc.) of the sample category, and the information of the corresponding sample table (such as the table name of the sample table, and the relevant information fields of the nodes for multi-node storage).

[0050] Correspondingly, after obtaining the sample table data in the data to be managed according to the data nature, save the sample table data as the first sample. The storage of the sample table data is stored in the form of a table, that is, the sample table data is stored in the form of a database table, and each sample table data corresponds to a sample category. Among them, the sample table data includes: sample ID, sample category identifier, sample collection time, location where the sample is collected, picture ID of the sample, number of annotations, sequence of annotation IDs, etc.

[0051] In this embodiment, the storing the second sample in the form of encoding is specifically as follows:

[0052] Input the sample position data into a preset grid model, so that the preset grid model outputs three-dimensional space grid data according to the sample position data; after binary conversion of the three-dimensional space grid data, generate a space grid code and store it.

[0053] Further, the preset grid model is a Beidou space map grid model. The establishment process of the Beidou space map grid model is as follows: perform data conversion on entity objects to obtain point, line, surface, and volume data, obtain a set of spatial grids, and based on the set of spatial grids, establish the Beidou space map grid model.

[0054] The step of inputting the sample position data into the preset grid model so that the preset grid model outputs three-dimensional spatial grid data according to the sample position data is specifically as follows: input the sample position data into the Beidou space map grid model, so that the preset grid model divides each large grid area in the sample position data into several small grid areas. After extracting the small grid areas to be counted occupied by spatial objects, perform data statistics on the small grid areas to be counted, and combine and generate spatial object interval data (i.e., three-dimensional spatial grid data) according to the statistical results and then output.

[0055] The step of performing binary conversion on the three-dimensional spatial grid data and then generating and storing a spatial grid code is specifically as follows: convert the spatial object interval data (i.e., three-dimensional spatial grid data) into a binary combination code, and according to the corresponding relationship of the large and small grid areas, convert the binary combination code into an absolute spatial grid code and store it.

[0056] In this embodiment, the third sample is stored in the form of a file, specifically as follows:

[0057] After obtaining the picture data and annotation result data in the data to be managed according to the data nature, use the picture data and the annotation result data as the third sample.

[0058] Storage of picture data: Store it in the form of directory + file. According to the size of the picture data volume, adopt a single-node or distributed storage strategy. When the data volume of a node has reached 20TB, store the newly incoming picture data in a new data node.

[0059] Storage of annotation result data: The annotation result data is stored in the form of directory + file. According to the size of the sample data volume, adopt a single-node or distributed storage strategy. When the data volume of a node has reached 20TB, store the newly incoming annotation result data in a new data node.

[0060] In this embodiment, it further includes:

[0061] Configure a table index strategy for the first sample, a grid index strategy for the second sample, and a directory file index strategy for the third sample.

[0062] In a specific embodiment, the step of configuring a table index strategy for the first sample is specifically as follows:

[0063] Since the first sample (i.e., sample category data and sample table information data) is stored in the form of a database table, for the need to quickly retrieve sample categories according to conditions such as road visual traffic conditions, sample category names, and sample information, the system speeds up the retrieval speed of sample categories through optimization methods such as adding data table indexes (i.e., table index strategies).

[0064] In a specific embodiment, the configuration of the grid index strategy for the second sample is specifically as follows:

[0065] Since the second sample is stored by encoding, and the second sample is sample position data, including the nature of time and space. If only through optimization means such as database tables for indexing in the spatial range, it is impossible to achieve fast retrieval of the second sample. Therefore, the grid space data retrieval method (i.e., grid index strategy) is used for indexing time and space for retrieval.

[0066] The grid space data retrieval method is specifically as follows: By obtaining the longitude and latitude of the point to be retrieved, performing an oblique axis equidistant azimuth projection on the longitude and latitude to obtain the spatial coordinate system coordinates of the point to be retrieved. According to the preset grid space that has been divided and has a coding rule and the spatial coordinate system coordinates of the point to be retrieved, all hierarchical spatial grids of the point to be retrieved in the grid space are obtained, all spatial objects within the spatial grid are extracted, the spatial object containing the retrieval point is extracted from all spatial objects, and the spatial object containing the point to be retrieved is used as the retrieval result.

[0067] In a specific embodiment, the configuration of the directory file index strategy for the third sample is specifically as follows:

[0068] Since the third sample is stored in the form of directory + file, when the storage size is greater than the limit, it is stored in multiple nodes. Therefore, the retrieval of picture data and annotation result data uses the "directory + file" method of the file system (i.e., directory file index strategy).

[0069] The retrieval process of the third sample is as follows:

[0070] 1. Input sample retrieval conditions, including query conditions such as sample category, sample time, and the position range of the point to be retrieved.

[0071] 2. Query the sample table name through the database query method according to the input sample category condition.

[0072] 3. Use the grid space data retrieval method to perform time and space retrieval on the sample data to obtain sample data that meets the conditions.

[0073] 4. Obtain the nodes where the samples and annotation results are located, the names and paths of the pictures, and the sequence IDs of the annotation results according to the qualified sample table data retrieved from the query, so as to obtain the sample picture data and the annotation results.

[0074] To further illustrate the sample data management device, please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a sample data management device provided by an embodiment of the present invention, including: an acquisition module 201 and a management module 202;

[0075] Among them, the acquisition module 201 is used to divide the to-be-managed sample data into a first sample, a second sample, and a third sample according to the data nature after acquiring the to-be-managed sample data;

[0076] The management module 202 is used to store the first sample in the form of a table, store the second sample in the form of encoding, and store the third sample in the form of a file.

[0077] In this embodiment, it further includes:

[0078] Configure a table index strategy for the first sample, configure a grid index strategy for the second sample, and configure a directory file index strategy for the third sample.

[0079] In this embodiment, the step of dividing the to-be-managed sample data into a first sample, a second sample, and a third sample according to the data nature is specifically as follows:

[0080] After obtaining the sample category data and the sample table information data in the to-be-managed data according to the data nature, use the sample category data and the sample table information data as the first sample;

[0081] After obtaining the sample position data in the to-be-managed data according to the data nature, use the sample position data as the second sample;

[0082] After obtaining the picture data and the annotation result data in the to-be-managed data according to the data nature, use the picture data and the annotation result data as the third sample.

[0083] In this embodiment, the step of storing the second sample in the form of encoding is specifically as follows:

[0084] Input the sample position data into a preset grid model, so that the preset grid model outputs three-dimensional space grid data according to the sample position data;

[0085] After performing binary conversion on the three-dimensional space grid data, generate a space grid code and store it.

[0086] An embodiment of the present invention provides a mobile terminal, including a processor and a memory. The memory stores computer-readable program code. When the processor executes the computer-readable program code, the steps of the above-mentioned sample data management method are implemented.

[0087] An embodiment of the present invention provides a storage medium that stores computer-readable program code. When the computer-readable program code is executed, the steps of the above-mentioned sample data management method are implemented.

[0088] In an embodiment of the present invention, after the acquisition module acquires the sample data to be managed, the sample data to be managed is divided into a first sample, a second sample, and a third sample according to the data nature; then, the management module stores the first sample in a tabular manner, stores the second sample in an encoded manner, and stores the third sample in a file manner.

[0089] In an embodiment of the present invention, different sample data is stored in different ways according to the data nature, which can flexibly achieve the effective storage of sample data. Retrieving different sample data according to different storage methods can effectively improve the retrieval efficiency of sample data.

[0090] Furthermore, in an embodiment of the present invention, the second sample is sample position data, and the third sample is picture data and annotation result data. Storing the sample position data in an encoded manner, that is, performing grid encoding and encoding compression on the sample position data, can reduce the storage of the sample position data. Storing the picture data and annotation result data in a file manner and storing the picture data and annotation result data in a file server in the directory + file manner can reduce the occupation of the database storage space.

[0091] Moreover, in an embodiment of the present invention, different indexing strategies are configured for different nature sample data, which can further improve the retrieval efficiency. For example: for the retrieval of the second sample, since the second sample is stored in an encoded manner, the grid space data retrieval method (i.e., the grid indexing strategy) is used for retrieval, and the retrieval result can be obtained quickly.

[0092] Finally, in an embodiment of the present invention, the sample data is managed by automatic configuration. Specifically, a single-node or distributed method is used for sample data management. When the data volume of a single node has reached 20TB, the new data is automatically stored in the next node, and there is no need for manual file migration and maintenance, solving the problem of regular maintenance of the server and high maintenance cost in the prior art.

[0093] After dividing the sample data to be managed into different samples according to the data nature in the embodiments of the present invention, the corresponding storage of different samples is respectively performed in the form of a table, in the form of coding, and in the form of a file.

[0094] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can be made, and these improvements and modifications are also regarded as the protection scope of the present invention.

Claims

1. A method for managing sample data, characterized in that, it includes: After obtaining the sample data to be managed, divide the sample data to be managed into a first sample, a second sample, and a third sample according to the data nature; The specific method of dividing the sample data to be managed into a first sample, a second sample, and a third sample according to the data nature is: After obtaining the sample category data and the sample table information data in the sample data to be managed according to the data nature, use the sample category data and the sample table information data as the first sample; After obtaining the sample position data in the sample data to be managed according to the data nature, use the sample position data as the second sample; After obtaining the picture data and the annotation result data in the sample data to be managed according to the data nature, use the picture data and the annotation result data as the third sample; Store the first sample in the form of a table, store the second sample in the form of encoding, and store the third sample in the form of a file; The specific method of storing the second sample in the form of encoding is: Input the sample position data into a preset grid model, so that the preset grid model outputs three-dimensional space grid data according to the sample position data; After performing binary conversion on the three-dimensional space grid data, generate a space grid code and store it.

2. A method for managing sample data according to claim 1, characterized in that, it further includes: Configure a table index strategy for the first sample, configure a grid index strategy for the second sample, and configure a directory file index strategy for the third sample.

3. A sample data management device, characterized in that, it includes: An acquisition module and a management module; Among them, the acquisition module is used to obtain the sample data to be managed, and divide the sample data to be managed into a first sample, a second sample, and a third sample according to the data nature; The specific method of dividing the sample data to be managed into a first sample, a second sample, and a third sample according to the data nature is: After obtaining the sample category data and the sample table information data in the sample data to be managed according to the data nature, use the sample category data and the sample table information data as the first sample; After obtaining the sample position data in the sample data to be managed according to the data nature, use the sample position data as the second sample; After obtaining the picture data and the annotation result data in the sample data to be managed according to the data nature, use the picture data and the annotation result data as the third sample; The management module is used to store the first sample in the form of a table, store the second sample in the form of encoding, and store the third sample in the form of a file; The specific method of storing the second sample in the form of encoding is: Input the sample position data into a preset grid model, so that the preset grid model outputs three-dimensional space grid data according to the sample position data; After performing binary conversion on the three-dimensional space grid data, generate a space grid code and store it.

4. A sample data management device according to claim 3, wherein, further comprising: configuring an index policy for the first sample configuration table, an index policy for the second sample configuration grid, and an index policy for the third sample configuration directory file.

5. A mobile terminal, wherein, comprising a processor and a memory, the memory storing computer-readable program code, and the processor implementing the steps of a sample data management method according to any one of claims 1 to 2 when executing the computer-readable program code.

6. A storage medium, wherein, the storage medium stores computer-readable program code, and when the computer-readable program code is executed, the steps of a sample data management method according to any one of claims 1 to 2 are implemented.

Citation Information

Patent Citations

  • Equipment and method for collecting picture samples

    CN109147093A

  • Information deduplication processing method and device and computer readable storage medium

    CN112749131A