Method and system for managing mass audios and audio annotation data
By formulating standardized verification specifications and establishing the relationship between audio and labeling files, the problems of separation of audio files and labeling information management, coarse granularity of version control and inefficient search efficiency are solved, and the efficient management of massive audio data and version traceability are achieved.
Patent Information
- Application Number
- CN202510222674.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-17
AI Technical Summary
When managing massive audio data in the prior art, the management of audio files and label information is separated, the version control is coarsely granular, and the search efficiency is inefficient, making it difficult to meet the compound needs of data quality, version traceability and efficient retrieval in training scenarios.
By formulating standardized verification specifications for audio and audio labeling files, organizing and uploading data, establishing the association relationship between audio and labeling files, supporting fine-grained editing and version management, and generating index directories based on search fields.
It realizes independent management and efficient retrieval of audio labeling information and audio files, supports fine-grained version control and multi-version management, and improves the flexibility and efficiency of data management.
Smart Images

Figure CN120164491A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data management, and in particular, to a method and system for managing massive audio and audio annotation data. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, the training requirements for audio-related models such as speech recognition and speech synthesis have increased sharply. The training of such models depends on massive audio files and their corresponding annotation data. The management of audio data involves a close association between the audio files themselves and the annotation information, and the annotation information usually needs to be iteratively modified frequently to adapt to different training requirements.
[0003] Currently, common file management solutions mainly include two categories: one is a storage system for general files, such as a cloud storage system, which supports file-level add, delete, and modify operations, but cannot independently manage the description information inside the file or associated with it, such as audio annotation text, nor can it record the fine-grained version changes of the annotation information; the other is a tool based on code version management, such as Git, which can achieve fine-grained version control of text content and multi-person collaboration, but its original design is for the scenario of small-size and high-frequency modification of code files, and it is difficult to adapt to the storage requirements of large-size audio files. In addition, existing tools usually treat the file and its description information as a whole for processing, lacking an independent management and associated retrieval mechanism for the two.
[0004] The core defect of the above two types of tools lies in the contradiction between their generality and domain adaptability. The file storage system, in pursuit of universality, does not design an independent management mechanism for the characteristics of audio data, resulting in the inability to decouple and modify the annotation information from the audio file, and lacking a fine-grained version recording function. Although the code management tool supports fine-grained version management, its architecture cannot support the efficient storage and retrieval of massive audio files. Ultimately, due to the lack of a unified management framework for the characteristics of audio data in the prior art, the following problems occur: the management of audio files and annotation information is fragmented, the version iteration record is coarse-grained, the retrieval efficiency is low, and it is difficult to meet the composite requirements of data quality, version traceability, and efficient retrieval in the training scenario. This problem has become a key bottleneck restricting the efficiency and standardization of audio data management. Summary of the Invention
[0005] This application provides a method and system for managing massive audio and audio annotation data, which solves the key problems such as the fragmentation of audio file and annotation information management, coarse-grained version control, and low retrieval efficiency in the prior art. This application provides the following technical solutions:
[0006] In a first aspect, this application provides a method for managing massive audio and audio annotation data, the method including:
[0007] Formulate a standardized verification specification for audio and audio annotation files, and organize the target audio and audio annotation files based on the standardized verification specification;
[0008] Upload the organized target audio and audio annotation files and verify them;
[0009] Establish an association relationship for the uploaded target audio and audio annotation files;
[0010] In response to a modification request initiated by the administrator, perform an editing operation on the uploaded target audio and audio annotation files;
[0011] Retrieve the target audio and audio annotation files based on the retrieval fields and generate an index directory.
[0012] In a specific implementable solution, the standardized verification specification includes audio format requirements, audio annotation fields, and annotation field requirements;
[0013] The audio format requirements support mainstream audio file formats;
[0014] The audio annotation fields include, but are not limited to, audio name, audio ID, audio text information, language, language variety, dialect, speech rate, acoustic environment, volume, and speaker attributes;
[0015] The annotation field requirements include, but are not limited to, field name, English identifier, field type, whether it is a required item, and field input requirements.
[0016] In a specific implementable solution, the organizing the target audio and audio annotation files based on the standardized verification specification includes:
[0017] Store the audio files and the corresponding annotation files in separate folders respectively, and establish a mapping association through the audio ID or audio name.
[0018] In a specific implementable solution, the uploading the organized target audio and audio annotation files and verifying them includes:
[0019] Upload the standard audio and audio annotation files through a script upload tool;
[0020] Based on the standardized verification specification, perform audio format verification, annotation field verification, and data correlation verification on the target audio and audio annotation files.
[0021] In a specific implementable solution, the establishing an association relationship for the uploaded target audio and audio annotation files includes:
[0022] Parse the uploaded target audio and audio annotation files, and match the audio files with their corresponding audio annotation files according to a predefined unique identifier;
[0023] Record the physical storage paths and relevant metadata of each file, and generate a mapping table in the database.
[0024] In a specific feasible implementation, the editing operation on the uploaded target audio and audio annotation files in response to a modification request initiated by a manager includes:
[0025] The editing operation includes, but is not limited to, adding, deleting, and editing;
[0026] After the editing operation is completed, a version record is generated, allowing the manager to record the version number;
[0027] Multiple version records are supported for the same file, and version number switching, version list viewing, and version comparison are supported.
[0028] In a specific feasible implementation, the retrieving the target audio and audio annotation files based on retrieval fields and generating an index directory includes:
[0029] The retrieval fields include text types and tag types;
[0030] The text type represents the text information of the audio, that is, the audio text of the audio;
[0031] The tag type represents various tag information of the audio.
[0032] In a second aspect, the present application provides a system for managing a large amount of audio and audio annotation data, adopting the following technical solutions:
[0033] A system for managing a large amount of audio and audio annotation data, including:
[0034] A specification formulation module, configured to formulate a standardized verification specification for audio and audio annotation files, and sort out the target audio and audio annotation files based on the standardized verification specification;
[0035] An association establishment module, configured to upload the sorted target audio and audio annotation files and verify them;
[0036] An upload verification module, configured to establish an association relationship for the uploaded target audio and audio annotation files;
[0037] An editing operation module, configured to perform an editing operation on the uploaded target audio and audio annotation files in response to a modification request initiated by a manager;
[0038] An index directory generation module, configured to retrieve the target audio and audio annotation files based on retrieval fields and generate an index directory.
[0039] In a third aspect, the present application provides an electronic device, which includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a method for managing a large amount of audio and audio annotation data as described in the first aspect.
[0040] In a fourth aspect, the present application provides a computer-readable storage medium, in which a program is stored, and when the program is executed by a processor, it is used to implement a method for managing a large amount of audio and audio annotation data as described in the first aspect.
[0041] In summary, the beneficial effects of the present application at least include:
[0042] 1) By establishing an independent audio annotation management mechanism, the annotation information can be stored and managed separately from the audio file, and independent retrieval is supported. Traditional audio management methods usually store the audio file and its annotation information as a whole, resulting in additional parsing and matching operations when querying or adjusting the annotation information, reducing the flexibility of data management. In the present application, a mapping relationship is established through the audio ID or audio name, enabling users to retrieve the annotation information separately, such as filtering by tags such as language, speech rate, or speaker attributes, and quickly locating the corresponding audio file after retrieving the target annotation information, realizing the efficient and independent management of the annotation data while maintaining its close association with the audio data.
[0043] 2) Support for field-level editing and modification functions, and provide version management and version push capabilities after data changes to ensure the complete traceability of information. Managers can modify some fields of the audio annotation information without having to re-upload the entire file. For example, when it is necessary to update the age or emotion label of the speaker of a certain audio, only this field can be modified without affecting other annotation data. Each modification will be recorded by the system, generating a unique version number, and storing the modifier, modification time, and change content to ensure that all adjustment processes can be traced and restored. This method not only improves the flexibility of data maintenance, but also effectively reduces the storage and management costs brought by repeated uploads, while improving the accuracy and consistency of the annotation data.
[0044] 3) It further provides the editability and version management functions for the audio file description information, enabling users to dynamically adjust the annotation information of the audio while maintaining the complete records of all historical versions. The traditional audio management methods usually only support static annotation and cannot record the historical changes of the annotation information, resulting in unclear data versions and difficulty in traceability. Through fine-grained version control, users can switch, compare, and roll back between different versions, ensuring data traceability and providing multi-version annotation data selection for different training needs. This ability is crucial for building high-quality and reproducible training datasets. Especially in application scenarios such as speech recognition and speech synthesis that require long-term iterative optimization, it can effectively improve data management efficiency and reduce maintenance costs.
[0045] By formulating standardized verification specifications, it realizes the unified format definition of audio files and their annotation information. A dedicated upload script tool is used to batch upload the sorted data. During the upload process, multiple verifications such as audio format, annotation field integrity, and data relevance are automatically executed to ensure that the uploaded data fully complies with the preset standards, providing high-quality data guarantee for subsequent storage and application. Then, the system parses all uploaded files, establishes the mapping relationship between audio and annotation files based on the predefined unique identifier, realizing the unified association and efficient traceability of data. It supports fine-grained editing of the uploaded data, constructs a multi-dimensional retrieval mechanism based on standardized fields, uses text information and label information to perform compound condition retrieval on audio and annotation data, and uniformly exports the qualified data as an index directory, facilitating directly calling the data in the training cluster for feature extraction or model training. At the same time, it supports users to modify and expand the directory content according to actual needs.
[0046] The above description is only an overview of the technical solution of this application. In order to be able to more clearly understand the technical means of this application and implement it in accordance with the content of the specification, the following uses the preferred embodiments of this application and combines with the attached drawings to elaborate in detail as follows. Brief Description of the Drawings
[0047] Figure 1 It is a schematic flowchart of the method for managing a large amount of audio and audio annotation data in the embodiment of this application.
[0048] Figure 2 It is a schematic overall flowchart of the method for managing a large amount of audio and audio annotation data in the embodiment of this application.
[0049] Figure 3 It is a structural block diagram of the system for managing a large amount of audio and audio annotation data in the embodiment of this application.
[0050] Figure 4 It is a block diagram of the electronic device for managing a large amount of audio and audio annotation data in the embodiment of this application. Specific Embodiments
[0051] The following will further describe in detail the specific embodiments of the present application in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present application, but are not used to limit the scope of the present application.
[0052] Optionally, the present application takes the method for managing massive audio and audio annotation data provided in each embodiment as an example for illustration in an electronic device. The electronic device is a terminal or a server. The terminal can be a mobile phone, a computer, a tablet computer, etc. The type of the electronic device is not limited in this embodiment.
[0053] Referring to Figure 1 , which is a schematic flowchart of a method for managing massive audio and audio annotation data provided in an embodiment of the present application. The method at least includes the following steps:
[0054] Step S101: Formulate a standardized verification specification for audio and audio annotation files, and organize the target audio and audio annotation files based on the standardized verification specification.
[0055] In step S101, in order to achieve the unity and standardization of data management, it is first necessary to formulate a set of standardized verification specifications for audio and audio annotation files. The specification includes three parts: audio format requirements, audio annotation fields, and annotation field requirements. Among them, the audio format requirements clearly stipulate the supported mainstream audio file formats, such as WAV, MP3, FLAC, etc. The audio annotation fields usually include information such as audio name, audio ID, audio text information, language, language variety, dialect, speech rate, acoustic environment, volume, and speaker attributes (such as gender, emotion, age). The annotation field requirements further clarify the field name, English identifier, field type, whether it is a required item, and field input requirements for each field.
[0056] After the standardized verification specification is formulated, organize the target audio and audio annotation files based on the standardized verification specification, that is, store the audio files and the corresponding annotation files in separate folders respectively, and establish a mapping association through the audio ID or audio name, so as to provide a unified, accurate, and reliable data basis for subsequent upload, automatic verification, storage, and retrieval.
[0057] Step S102: Upload the organized target audio and audio annotation files and verify them.
[0058] In implementation, the standard audio and audio annotation files are uploaded through a script upload tool, and automated verification is performed during the upload process to verify the consistency of the data based on the standardized verification specifications. The verification content includes but is not limited to the following aspects: First, audio format verification, checking whether the audio file meets the format requirements defined in step S101, such as whether it is a supported format, and whether the sampling rate, bit depth, and number of channels meet the specifications. Second, annotation field verification, ensuring the field integrity of all annotation information, checking whether there are missing required fields, whether the field formats are correct, and whether the text fields exceed the length limit. Third, data correlation verification, verifying the matching relationship between the audio file and the annotation file, ensuring that the audio ID or audio name is consistent with the annotation file, and avoiding the situation of missing files or incorrect association of annotation information.
[0059] In addition, the method also provides an error feedback mechanism. For the data that fails the verification, the system will automatically reject it and give a detailed error prompt for re-uploading after correction. Through the above automated verification, the standardization of the data in the database can be further improved, providing high-quality data guarantee for subsequent applications such as storage, retrieval, and model training.
[0060] Step S103: Establish an association relationship for the uploaded target audio and audio annotation files.
[0061] In step S103, an association relationship will be established for the uploaded target audio and audio annotation files. Specifically, first, all uploaded files are parsed. Based on the predefined unique identifier (such as the audio ID or audio name), the audio file is accurately matched with its corresponding audio annotation file. At the same time, the physical storage paths and relevant metadata of each file are recorded, and an "audio-annotation-storage location" mapping table is generated in the database to achieve unified management and efficient traceability of the data.
[0062] Step S104: Respond to the modification request initiated by the administrator and perform editing operations on the uploaded target audio and audio annotation files.
[0063] In step S104, in response to the modification request initiated by the administrator, the administrator is supported to perform editing operations on the uploaded target audio and audio annotation files, including operations such as adding, deleting, and editing.
[0064] In implementation, authorized managers can perform fine-grained editing on the uploaded data according to actual needs, such as adding new files, deleting redundant data, or making partial modifications to the file content. After each editing operation is completed, a version record will be automatically generated, and managers are allowed to assign a version number to this batch of data. The version information recorded includes detailed information such as the modifier, a summary of the modification content, and the modification time. Multiple version records are supported for the same file, and basic functions such as version number switching, version list viewing, and version comparison are available to ensure that the historical change process of the data is completely preserved and traceable. The above version management mechanism not only facilitates subsequent problem troubleshooting and data tracing but also provides strong guarantee for the reproduction of training data.
[0065] Step S105: Retrieve the target audio and audio annotation files based on the retrieval fields and generate an index directory.
[0066] Specifically, the retrieval fields include text-based and label-based ones. The text-based ones are the text information of the audio, usually the audio text of this segment of audio. The label-based ones are various label information of the audio, such as speech rate, speaker age, etc. After retrieving according to the retrieval fields, the qualified audio data and corresponding annotation information are uniformly exported, and an index directory containing the physical storage address of each audio file, relevant annotation data, and necessary metadata is generated. This directory does not actually store the physical audio but provides a path pointing to the audio storage location, facilitating direct invocation in the training cluster for feature extraction or model training. At the same time, it supports users to modify and expand the directory content according to actual needs, thus realizing the efficient decoupling and precise control of the data management and application processes.
[0067] In summary, in combination with Figure 2, through the establishment of a complete process including standardized verification specifications, data collation, upload verification, automatic association, edit version management, and efficient retrieval index generation, the key problems in the prior art such as the fragmentation of audio file and annotation information management, coarse-grained version control, and low retrieval efficiency are solved. First, by formulating a set of standardized verification specifications including audio format requirements, audio annotation field definitions, and annotation field requirements, the unified format definition of audio files and their annotation information is realized, and based on this specification, the target audio and annotation files are stored in separate folders respectively, and the mapping association is established through the audio ID or audio name, thus providing a unified and accurate data basis for subsequent operations. Subsequently, a dedicated upload script tool is used to batch upload the collated data, and multiple verifications such as audio format, annotation field integrity, and data relevance are automatically executed during the upload process to ensure that the uploaded data fully complies with the preset standards and provide high-quality data guarantee for subsequent storage and application. Then, the system parses all uploaded files, establishes the mapping relationship between audio and annotation files based on the predefined unique identifier, realizes the unified association and efficient traceability of data. When responding to the modification request of the manager, the system supports fine-grained editing of the uploaded data, including adding new files, deleting redundant data, or making partial modifications to the file content, and automatically generates a detailed version record after each edit operation, recording the modifier, modification content, and modification time, realizing refined version management, version switching, and historical traceability to meet the needs of continuous iterative update of training data. Finally, the system constructs a multi-dimensional retrieval mechanism based on standardized fields, uses text information and label information to perform compound condition retrieval on audio and annotation data, and exports the qualified data as an index directory uniformly. This directory details the physical storage address, associated annotation data, and necessary metadata of each audio file, and provides a path pointing to the actual storage location, facilitating the direct invocation of data in the training cluster for feature extraction or model training, and at the same time supporting users to modify and expand the directory content according to actual needs. Through the above process, this application realizes the decoupled management, fine version tracking, and efficient retrieval application of audio files and annotation information, thus effectively solving the problems of unrefined version, low retrieval efficiency, and data inconsistency in the prior art in the management of massive audio data, and providing high-quality, traceable, and easy-to-manage data support for audio model training such as speech recognition and speech synthesis.
[0068] Figure 3 FIG. is a structural block diagram of a system for managing massive audio and audio annotation data provided by an embodiment of this application. The system at least includes the following modules:
[0069] Specification formulation module, used to formulate standardized verification specifications for audio and audio annotation files, and collate the target audio and audio annotation files based on the standardized verification specifications;
[0070] An upload verification module, configured to upload the organized target audio and audio annotation files and verify them;
[0071] An association establishment module, configured to establish an association relationship for the uploaded target audio and audio annotation files;
[0072] An editing operation module, configured to perform an editing operation on the uploaded target audio and audio annotation files in response to a modification request initiated by a manager;
[0073] An index directory generation module, configured to retrieve the target audio and audio annotation files based on retrieval fields and generate an index directory.
[0074] For relevant details, refer to the above method embodiments.
[0075] Figure 4 It is a block diagram of an electronic device provided by an embodiment of the present application. The device at least includes a processor 401 and a memory 402.
[0076] The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 401 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 401 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0077] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include high-speed random access memory, as well as non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 is used to store at least one instruction for being executed by the processor 401 to implement the method for managing massive audio and audio annotation data provided in the method embodiments of the present application.
[0078] In some embodiments, the electronic device may further optionally include: a peripheral device interface and at least one peripheral device. The processor 401, the memory 402, and the peripheral device interface may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface through a bus, signal lines, or a circuit board. Schematically, the peripheral devices include, but are not limited to: a radio frequency circuit, a touch display screen, an audio circuit, and a power supply, etc.
[0079] Certainly, the electronic device may also include fewer or more components, and this embodiment does not limit this.
[0080] Optionally, the present application also provides a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the method for managing massive audio and audio annotation data in the above method embodiments.
[0081] Optionally, the present application also provides a computer product, which includes a computer-readable storage medium, in which a program is stored, and the program is loaded and executed by a processor to implement the method for managing massive audio and audio annotation data in the above method embodiments.
[0082] The technical features of the above embodiments may be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0083] The above embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several deformations and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.
Claims
1. A method for managing massive audio and audio annotation data, characterized in that: The method comprises: Formulate standardized verification specifications for audio and audio annotation files, and organize the target audio and audio annotation files based on the standardized verification specifications; Upload the organized target audio and audio annotation files and verify them; Establishing an association relationship between the uploaded target audio and the audio annotation file; In response to a modification request initiated by an administrator, editing the uploaded target audio and audio annotation file; The target audio and audio annotation files are retrieved based on the search fields and an index directory is generated.
2. The method for managing massive audio and audio annotation data according to claim 1, characterized in that: The standardized verification specifications include audio format requirements, audio annotation fields, and annotation field requirements; The audio format is required to support mainstream audio file formats; The audio annotation fields include but are not limited to audio name, audio ID, audio text information, language, language, dialect, speaking speed, acoustic environment, volume and speaker attributes; The marking field requirements include but are not limited to field name, English label, field type, whether it is a required item and field input requirements.
3. The method for managing massive audio and audio annotation data according to claim 1, characterized in that: The arranging the target audio and the audio annotation file based on the standardized verification specification includes: Store the audio files and the corresponding annotation files in separate folders, and establish a mapping association through the audio ID or audio name.
4. The method for managing massive audio and audio annotation data according to claim 1, characterized in that: The uploading and verifying of the organized target audio and audio annotation files includes: Upload the marked audio and audio annotation files through the script upload tool; The target audio and audio annotation files are subjected to audio format verification, annotation field verification, and data relevance verification based on standardized verification specifications.
5. The method for managing massive audio and audio annotation data according to claim 1, characterized in that: The establishing of an association relationship between the uploaded target audio and the audio annotation file includes: Parsing the uploaded target audio and audio annotation file, and matching the audio file with its corresponding audio annotation file according to a predefined unique identifier; Record the physical storage path and related metadata of each file, and generate a mapping table in the database.
6. The method for managing massive audio and audio annotation data according to claim 1, characterized in that: In response to the modification request initiated by the administrator, editing the uploaded target audio and audio annotation file includes: The editing operations include but are not limited to adding, deleting, and editing; A version record is generated after the editing operation is completed, allowing the administrator to record the version number; The same file supports multiple version records, version number switching, version list viewing and version comparison.
7. The method for managing massive audio and audio annotation data according to claim 1, characterized in that: The step of retrieving the target audio and audio annotation files based on the search field and generating an index directory comprises: The search fields include text and label types; The text class represents the text information of the audio, that is, the audio text of the audio; The tag class represents various types of tag information of audio.
8. A system for managing massive audio and audio annotation data, characterized in that: include: A specification formulation module, used to formulate standardized verification specifications for audio and audio annotation files, and to organize the target audio and audio annotation files based on the standardized verification specifications; An association establishment module, used to upload the sorted target audio and audio annotation files and verify them; An upload verification module, used to establish an association relationship between the uploaded target audio and the audio annotation file; An editing operation module, used to edit the uploaded target audio and audio annotation files in response to a modification request initiated by an administrator; The index directory generating module is used to retrieve the target audio and audio annotation files based on the search field and generate an index directory.
9. An electronic device, characterized in that: The device includes a processor and a memory; a program is stored in the memory, and the program is loaded and executed by the processor to implement a method for managing massive audio and audio annotation data as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The storage medium stores a program, and when the program is executed by the processor, it is used to implement a method for managing massive audio and audio annotation data according to any one of claims 1 to 7.