Multi-modal dynamic labeling method and system based on micro-kernel architecture, electronic device and storage medium
Patent Information
- Application Number
- CN202610986010.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-29
AI Technical Summary
[0012]核心技术问题:多模态数据标注工具碎片化、无法统一管理的问题
[0061]本发明通过“空壳框架+功能模板”的设计,实现了“一套系统,多种标注”。用户通过勾选配置即可生成定制化标注界面,极大降低了开发和维护成本;
Smart Images

Figure CN122837812A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of vehicle-mounted intelligent agents, and more particularly to a multimodal dynamic annotation method based on a microkernel architecture, a multimodal dynamic annotation system based on a microkernel architecture, electronic devices, and storage media. Background Technology
[0002] With the development of artificial intelligence technology, data annotation, as the "fuel" for model training, is becoming increasingly important. Currently, the mainstream annotation tools on the market can be mainly divided into two categories:
[0003] General-purpose open-source tools, such as LabelMe and LabelImg, are mainly used for single image object detection tasks. Their functions are fixed and their scalability is poor.
[0004] Vertical domain SaaS platforms, such as Prodigy for NLP or Audacity for speech (with scripts), are typically designed for specific data types and have a high degree of coupling between interface and functionality.
[0005] In the field of automotive intelligence (such as autonomous driving and smart cockpits), it is often necessary to process multiple modalities of data simultaneously, including images (road condition recognition), audio (voice commands), and text (user intent). Existing technologies have the following significant shortcomings:
[0006] Problems with existing technology:
[0007] The tools are severely fragmented: enterprises need to purchase or maintain multiple independent annotation systems (one for map annotation, one for voice annotation, and one for text annotation), resulting in data dispersion, high management costs, and difficulties in collaboration.
[0008] Poor functionality reusability: The UI components of the existing annotation platform are strongly bound to the business logic. If a new annotation type is to be added (such as expanding from "single-turn dialogue" to "multi-turn dialogue"), the front-end interface often needs to be refactored, resulting in a long development cycle.
[0009] Low level of intelligence and limited efficiency improvement methods: Although some platforms have introduced AI pre-annotation, it is usually a hard-coded model for specific tasks. For example, speech-to-text (ASR) models and text entity extraction models cannot be flexibly switched on the same interface, making it difficult to support multimodal fusion annotation scenarios.
[0010] Lack of unified task orchestration capabilities: Annotation administrators cannot flexibly define task flows for mixed "audio + text" input through a unified "task configuration center". Summary of the Invention
[0011] The purpose of this invention is to provide a microkernel-based multimodal dynamic annotation method, a microkernel-based multimodal dynamic annotation system, an electronic device, and a storage medium, thereby solving at least one of a number of technical problems.
[0012] Core technical challenges include: fragmented multimodal data annotation tools and the inability to manage them uniformly; difficulties in expanding annotation functionality and the need for repeated UI development; and the inability to dynamically match AI pre-annotation models with annotation tasks, resulting in minimal efficiency improvements.
[0013] This invention provides the following solution:
[0014] According to a first aspect of the present invention, a multimodal dynamic annotation method based on a microkernel architecture is provided, comprising:
[0015] S1. Receive task configuration parameters input from the administrator;
[0016] Task configuration parameters include: target data modality type and corresponding functional component identifier;
[0017] S2. Parse the task configuration parameters and extract the corresponding functional component identifiers;
[0018] Based on a preset plugin standard protocol, matching front-end atomic functional components are dynamically loaded; among them, the corresponding functional components are dynamically loaded according to the configured template ID.
[0019] S3. By using a pre-set empty shell framework, the loaded functional components are laid out and rendered to generate a customized annotation interface that adapts to the target data modality type.
[0020] S4. For the target data modality of the current annotation task, call the matching AI model service, obtain the AI pre-annotation results and inject them into the annotation interface;
[0021] S5. Manually correct the pre-annotation results, receive the corrected annotation data, and complete data verification and persistent storage.
[0022] Furthermore, including:
[0023] A basic empty shell framework with no business logic is used as an empty shell framework;
[0024] Based on an empty shell framework, it provides standard HTML container, event bus, layout management and data basic adaptation capabilities;
[0025] All labeling business functions are implemented through callable functional components.
[0026] Furthermore, including:
[0027] The default plugin standard protocol defines a unified IAnnotationPlugin interface lifecycle specification for each functional component;
[0028] The lifecycle specification includes: the init method for initialization, the render method for UI rendering, and the destroy method for component destruction;
[0029] Each functional component is deployed and dynamically loaded independently based on a lifecycle approach.
[0030] Furthermore, including:
[0031] Functional components communicate with each other via an event bus within the shell framework, using a publish-subscribe pattern.
[0032] In this way, component collaboration is achieved by subscribing to preset events, decoupling the direct coupling and calls between various functional components.
[0033] Furthermore, including:
[0034] The target data modality types include any one or more combinations of audio modality, text modality, and image / video modality;
[0035] Different functional components and AI models correspond to different modal configurations;
[0036] When the target data modality is audio, load the AudioPlayer audio playback component, the TimestampInput timestamp input component, and the TranscriptEditor rich text editing component;
[0037] The ASR speech recognition model is invoked to generate pre-annotated text results with timestamps, and words with confidence scores below a preset confidence threshold are highlighted to remind the human correction end to review them.
[0038] When the target data modality is text, load the QueryDisplay text display component, the Tagselector tag selection component, and the EntityRelationGraph entity relationship graph display component;
[0039] The NLP semantic understanding model is invoked to complete text intent recognition and automatic slot filling, and the recognized slots are highlighted with differentiated colors.
[0040] When the target data modality is an image / video modality, load the CanvasRenderer rendering engine, ShapeDrawer graphics drawing toolbar, ClassSelector object category label selection component, and TimelineScrubber video timeline component;
[0041] Call the CV vision detection model to generate target coordinates and category pre-labeled bounding boxes, which can be used to support manual drag-and-drop correction of boundaries and adjustment of category labels;
[0042] When the target data modality is a multimodal combination modality, the corresponding multiple functional components are loaded synchronously;
[0043] Multi-component timeline synchronization is achieved through an event bus, enabling multi-modal fusion annotation.
[0044] Furthermore, it also includes mechanisms for handling AI service anomalies and for circuit breaker fault tolerance:
[0045] Monitor the service call status of AI models, and trigger service degradation strategies when service timeouts, recognition confidence scores fall below a preset threshold, network anomalies occur, or a preset number of consecutive call failures occur.
[0046] Service degradation strategies include: activating AI-assisted functions via circuit breakers, switching to a purely manual annotation mode, and caching local data using IndexedDB.
[0047] Data will be automatically synchronized once the network is restored.
[0048] Furthermore, it also includes data consistency guarantee mechanisms:
[0049] The front-end annotation view is updated in real time using the OptimisticUI optimistic update method;
[0050] If backend data storage fails, the rollback view rollback method is triggered, restoring the view to its state before the annotation was edited and displaying an error message.
[0051] According to a second aspect of the present invention, a microkernel-based multimodal dynamic annotation system is provided for implementing a microkernel-based multimodal dynamic annotation method, comprising: an empty shell framework layer, a functional plugin pool, a backend service layer, an AI model service layer, and a data persistence layer;
[0052] The empty shell framework layer is used to provide basic layout rendering, event bus communication, task configuration parsing, user authentication and data persistence capabilities;
[0053] A functional plugin pool stores various atomic, pluggable annotation functional components, supporting dynamic loading on demand.
[0054] The backend service layer is used to parse task configurations, route and schedule AI models, and process data verification and storage requests.
[0055] The AI model service layer integrates ASR speech recognition model, NLP semantic understanding model, and CV visual detection model to output pre-annotated results with metadata for each modality.
[0056] The data persistence layer is used to store annotation task configurations and final annotation data.
[0057] According to a third aspect of the present invention, an electronic device is provided, comprising: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0058] The memory stores computer programs, which, when executed by the processor, cause the processor to perform steps such as those of a microkernel-based multimodal dynamic annotation method.
[0059] According to a fourth aspect of the present invention, a computer-readable storage medium is provided, comprising: storing a computer program executable by an electronic device, wherein when the computer program is run on the electronic device, the electronic device performs steps such as those of a microkernel-based multimodal dynamic annotation method.
[0060] The above solution achieves the following beneficial technical effects:
[0061] This invention achieves "one system, multiple annotations" through a "hollow framework + functional template" design. Users can generate customized annotation interfaces simply by selecting configuration options, greatly reducing development and maintenance costs.
[0062] This invention reduces the number of manual keyboard inputs and mouse clicks and improves annotation efficiency by accessing large models / semantic models for pre-identification / pre-annotation (such as automatically filling slots and generating ASR text).
[0063] The visual configuration function of this invention enables non-technical personnel to quickly build annotation projects that meet business needs;
[0064] The underlying player control, text editor, and large model API interface of this invention are all encapsulated into a general module that can be reused in different annotation modes. Attached Figure Description
[0065] Figure 1 This is a flowchart of a multimodal dynamic annotation method based on a microkernel architecture provided by one or more embodiments of the present invention.
[0066] Figure 2 This is a structural diagram of a multimodal dynamic annotation system based on a microkernel architecture provided by one or more embodiments of the present invention.
[0067] Figure 3 This is a schematic diagram of the overall system architecture provided in a specific embodiment of the present invention.
[0068] Figure 4This is a schematic diagram of annotation task configuration and interface rendering provided in a specific embodiment of the present invention.
[0069] Figure 5 This is a schematic diagram of a multimodal annotation interaction process provided in a specific embodiment of the present invention.
[0070] Figure 6 This is a schematic diagram of the plug-in core constraints and communication specifications provided in a specific embodiment of the present invention.
[0071] Figure 7 This is a schematic diagram of the timing of AI anomaly fallback and data consistency provided in a specific embodiment of the present invention.
[0072] Figure 8 This is a block diagram of an electronic device structure based on a microkernel architecture multimodal dynamic annotation method provided by one or more embodiments of the present invention. Detailed Implementation
[0073] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0074] Figure 1 This is a flowchart of a multimodal dynamic annotation method based on a microkernel architecture provided by one or more embodiments of the present invention.
[0075] Example 1, as Figure 1 The microkernel-based multimodal dynamic annotation method shown includes:
[0076] S1. Receive task configuration parameters input from the administrator;
[0077] Task configuration parameters include: target data modality type and corresponding functional component identifier;
[0078] S2. Parse the task configuration parameters and extract the corresponding functional component identifiers;
[0079] Based on a preset plugin standard protocol, matching front-end atomic functional components are dynamically loaded; wherein, according to the configured template ID, the corresponding functional components are dynamically loaded.
[0080] S3. By using a pre-set empty shell framework, the loaded functional components are laid out and rendered to generate a customized annotation interface that adapts to the target data modality type.
[0081] S4. For the target data modality of the current annotation task, call the matching AI model service, obtain the AI pre-annotation results and inject them into the annotation interface;
[0082] S5. Manually correct the pre-annotation results, receive the corrected annotation data, and complete data verification and persistent storage.
[0083] In this embodiment, it includes:
[0084] A basic empty shell framework with no business logic is used as an empty shell framework;
[0085] Based on an empty shell framework, it provides standard HTML container, event bus, layout management and data basic adaptation capabilities;
[0086] All labeling business functions are implemented through callable functional components.
[0087] In this embodiment, it includes:
[0088] The default plugin standard protocol defines a unified IAnnotationPlugin interface lifecycle specification for each functional component;
[0089] The lifecycle specification includes: the init method for initialization, the render method for UI rendering, and the destroy method for component destruction;
[0090] Each functional component is deployed and dynamically loaded independently based on a lifecycle approach.
[0091] In this embodiment, it includes:
[0092] Functional components communicate with each other via an event bus within the shell framework, using a publish-subscribe pattern.
[0093] In this way, component collaboration is achieved by subscribing to preset events, decoupling the direct coupling and calls between various functional components.
[0094] In this embodiment, it includes:
[0095] The target data modality types include any one or more combinations of audio modality, text modality, and image / video modality;
[0096] Different functional components and AI models correspond to different modal configurations;
[0097] When the target data modality is audio, load the AudioPlayer audio playback component, the TimestampInput timestamp input component, and the TranscriptEditor rich text editing component;
[0098] The ASR speech recognition model is invoked to generate pre-annotated text results with timestamps, and words with confidence scores below a preset confidence threshold are highlighted to remind the human correction end to review them.
[0099] When the target data modality is text, load the QueryDisplay text display component, the Tagselector tag selection component, and the EntityRelationGraph entity relationship graph display component;
[0100] The NLP semantic understanding model is invoked to complete text intent recognition and automatic slot filling, and the recognized slots are highlighted with differentiated colors.
[0101] When the target data modality is an image / video modality, load the CanvasRenderer rendering engine, ShapeDrawer graphics drawing toolbar, ClassSelector object category label selection component, and TimelineScrubber video timeline component;
[0102] Call the CV vision detection model to generate target coordinates and category pre-labeled bounding boxes, which can be used to support manual drag-and-drop correction of boundaries and adjustment of category labels;
[0103] When the target data modality is a multimodal combination modality, the corresponding multiple functional components are loaded synchronously;
[0104] Multi-component timeline synchronization is achieved through an event bus, enabling multi-modal fusion annotation.
[0105] This embodiment also includes mechanisms for AI service anomaly mitigation and circuit breaker fault tolerance:
[0106] Monitor the service call status of AI models, and trigger service degradation strategies when service timeouts, recognition confidence scores fall below a preset threshold, network anomalies occur, or a preset number of consecutive call failures occur.
[0107] Service degradation strategies include: activating AI-assisted functions via circuit breakers, switching to a purely manual annotation mode, and caching local data using IndexedDB.
[0108] Data will be automatically synchronized once the network is restored.
[0109] This embodiment also includes a data consistency guarantee mechanism:
[0110] The front-end annotation view is updated in real time using the OptimisticUI optimistic update method;
[0111] If backend data storage fails, the rollback view rollback method is triggered, restoring the view to its state before the annotation was edited and displaying an error message.
[0112] Figure 2 This is a structural diagram of a multimodal dynamic annotation system based on a microkernel architecture provided by one or more embodiments of the present invention.
[0113] Example 2, as Figure 2 The microkernel-based multimodal dynamic annotation system shown is used to implement the microkernel-based multimodal dynamic annotation method, including: a shell framework layer, a functional plugin pool, a backend service layer, an AI model service layer, and a data persistence layer;
[0114] The empty shell framework layer is used to provide basic layout rendering, event bus communication, task configuration parsing, user authentication and data persistence capabilities;
[0115] A functional plugin pool stores various atomic, pluggable annotation functional components, supporting dynamic loading on demand.
[0116] The backend service layer is used to parse task configurations, route and schedule AI models, and process data verification and storage requests.
[0117] The AI model service layer integrates ASR speech recognition model, NLP semantic understanding model, and CV visual detection model to output pre-annotated results with metadata for each modality.
[0118] The data persistence layer is used to store annotation task configurations and final annotation data.
[0119] It is worth noting that although this system / device only discloses the above-mentioned modules / units, it does not mean that this system / device is limited to the above-mentioned basic functional modules. On the contrary, what this invention intends to express is that, based on the above-mentioned basic functional modules, those skilled in the art can add one or more functional modules in combination with the prior art to form an infinite number of embodiments or technical solutions. That is to say, this system is open rather than closed. It cannot be assumed that the scope of protection of the claims of this invention is limited to the above-disclosed basic functional modules just because this embodiment only discloses a few basic functional modules.
[0120] Example 3 discloses a multi-mode annotation platform based on a microkernel architecture in one specific embodiment. The core of the platform is a shell framework responsible for layout rendering, user authentication, and data persistence; the periphery is a plugin pool containing various atomic functional components.
[0121] Example 4, Core Solution Flow:
[0122] Task configuration phase: Administrators create tasks, select data types (audio / text / video) and corresponding annotation templates;
[0123] UI rendering stage: The front-end framework dynamically loads the corresponding functional components (Widgets) based on the configured template ID.
[0124] AI-assisted phase: The backend service routes to the corresponding AI model (ASR / NLP model) based on the task type and returns the pre-labeled results to populate the frontend;
[0125] Data submission stage: Validate the data format and store it in the database.
[0126] Example 5, Specific Module Design:
[0127] 1) Design an empty shell frame (such as...) Figure 3 (As shown)
[0128] This framework does not contain specific business logic; it only provides a standard HTML container and an event bus. It defines the left side as the "data source display area," the middle as the "main annotation area," and the right side as the "attribute / tag configuration area."
[0129] The core focus is on the complete interaction flow of "receiving configuration instructions -> calling the empty framework -> dynamically loading the functional components of the corresponding modality -> rendering the annotation interface".
[0130] 2) Audio recognition content annotation template (e.g.) Figure 4 (As shown)
[0131] Includes the following functional components:
[0132] AudioPlayer: Audio waveform display pre-play control
[0133] TimestampInput: Timestamp input field
[0134] TranscriptEditor: A rich text editing box (used for ASR result correction).
[0135] The system automatically calls the ASR large model interface to preprocess the uploaded audio file, automatically fills the returned text results with timestamps into the TranscriptEditor, and highlights the words with the highest confidence level, prompting manual review.
[0136] 3) NLP Text Annotation Templates
[0137] Includes the following functional components:
[0138] QueryDisplay: A text query display box (supports highlighted selection)
[0139] Tagselector: Tag system dropdown / selection component
[0140] EntityRelationGraph: A graph display of entity relationships
[0141] The system calls an NLU semantic model to perform intent recognition and slot filling on the input text. For example, if the user enters "Help me open the car window," the model automatically recognizes the intent as "open window" and marks "car window" as a "carcontrol" entity. The front end automatically identifies the slots and highlights them with different colors. The user only needs to click to confirm or modify, without needing to manually change anything.
[0142] 4) Image / Video Object Detection and Classification Labeling Template
[0143] To achieve standardized labeling of visual data from autonomous driving and smart cockpits, the system activates the following visualization function components when the administrator selects the "image" or "video" data type.
[0144] CanvasRenderer: A canvas rendering engine that supports loading static images or loading video streams frame by frame.
[0145] ShapeDrawer: A toolbar for drawing geometric shapes (supports rectangles, polygons, and keypoints).
[0146] ClassSelector: Object category label selector (e.g., pedestrians, vehicles, traffic signs, lane lines).
[0147] TimelineScrubber: Video progress bar and timeline, supports forward / backward movement by frame.
[0148] The system calls upon large CV models (such as the YOLO series and SAM segmentation models) to automatically detect objects in the first frame of uploaded images or videos.
[0149] Automatic filling: The recognition results returned by the model (such as coordinates (x,y,z):[100,200,300], category: car) are automatically converted into pre-labeled bounding boxes on the canvas. Annotators do not need to annotate from scratch; they only need to check whether the position and category of the pre-labeled bounding boxes are correct, and adjust the boundaries or correct the category labels by dragging, thereby greatly reducing mouse clicks and dragging distances.
[0150] 5) Multimodal interaction annotation of intelligent cockpit (e.g.) Figure 5 (As shown)
[0151] Scenario description: The input data is a data stream containing the driver's voice (Audio), facial expression video (Video), and speech-to-text (Text).
[0152] Annotation process: Annotators see the video on the left (to determine distraction), hear the audio in the middle (to determine emotion), and directly associate the text slot on the right (to determine intent).
[0153] Architecture support: The empty framework loads three plugins at the same time: VideoPlayer, AudioPlayer and TextEditor, and realizes timeline synchronization through the event bus (such as dragging the video progress bar, and the audio and text highlighting are linked accordingly).
[0154] like Figure 6 As shown, the PluginProtocol defines the interface standard (such as the IAnnotationPlugin interface) that all functional components must implement, and includes three lifecycle methods: init() (initialization), render(container) (rendering to the container), and destroy (destruction).
[0155] Communication Specification: Introduces a publish / subscribe pattern (Pub / Sub). Plugins do not call each other directly; instead, they send messages through the EventBus of the shell framework. For example, when AudioPlayer finishes playback, it publishes an audio:ended event, and TranscriptEditor subscribes to this event to automatically lock the editing process.
[0156] Conflict handling and version management: rendering overlap: The framework has a built-in LayoutManager that uses a Z-Index stack to manage layers, or uses placeholders to force the grid layout to prevent plugins from arbitrarily occupying DOM nodes.
[0157] Version upgrade: The backend configuration table records the plugin version number (SemVer). When the frontend loads, it checks the compatibility matrix. If an incompatible plugin combination is detected, the administrator is notified and task creation is blocked.
[0158] In this process, if AI pre-identification fails, the AI failure fallback mechanism is as follows (e.g.) Figure 7 (as shown)
[0159] Add flowcharts or pseudocode descriptions during the AI-assisted phase:
[0160] Timeout handling: If the ASR service response takes more than 5 seconds, the front end will automatically switch to "manual entry mode" and the back end will retry asynchronously.
[0161] Recognition Failure: If the model returns a confidence score below the threshold (e.g., <0.3) or an error code, the system will automatically hide the pre-annotation box, display only the original audio / text, and display the message "AI pre-recognition failed, please annotate manually" on the interface.
[0162] Error rollback: When there is a network failure, IndexedDB is used for local caching to ensure that the annotator can still operate, and it will automatically synchronize after the network is restored.
[0163] Cross-layer communication: The DTO (Data Transfer Object) standard is used to specify the data format for each layer. For example, the data returned by the AI service layer must be converted into an AnnotationData structure that the rendering layer can recognize by the Adapter.
[0164] AI anomaly: The Circuit Breaker mode is adopted. If the AI model fails to be called 10 times consecutively, the system will automatically break the circuit, gray out the "AI Assist" button on the front-end interface, and only retain the basic annotation function to prevent the UI from freezing.
[0165] Data consistency: Introducing OptimisticUI and Rollback mechanisms. When a user clicks "Save", the front-end view is updated first, and a request is sent to the back-end. If the back-end fails to save (returns a 500 error), the front-end triggers the rollback() method, restoring to the previous state and displaying "Save failed, please try again".
[0166] Example 6: Annotation process for in-vehicle voice commands:
[0167] The administrator logs into the backend, creates a new task, and selects "audio annotation" as the task type.
[0168] On the template configuration page, check "Audio Player", "ASR Pre-recognition", and "Text Correction Box".
[0169] The annotator logs in to the front end and opens the task. The left side of the interface displays the audio waveform, and the right side displays a blank text input box.
[0170] The system backend asynchronously calls the ASR model, instantly converting the audio into text: "Navigate to the nearest gas station".
[0171] Upon hearing the audio playback, the annotator discovered that ASR had misidentified "gas station" as "travel station" and corrected it directly in the text box.
[0172] The annotator clicks submit, and the data is entered into the database.
[0173] Example 7: Intelligent Cockpit Multi-Turn Dialogue Intent Annotation:
[0174] The administrator creates a new task, selecting "text annotation" as the task type.
[0175] On the template configuration page, check "Query Display", "Intent Tag Tree" and "Slot Highlight", and import the preset tag system (such as: Navigation, Multimedia, Vehicle Control).
[0176] The annotator saw the text: "The weather is nice today, let's play a Jay Chou song."
[0177] The system calls the large NLP model and automatically highlights the text: [Context:Chat][Music:Play][Singer:JayChou].
[0178] The labeler only needs to check if the label is correct and click save.
[0179] Example 8: A dynamic interface generation method based on plug-in configuration: receiving task configuration parameters, parsing the functional component identifiers contained in the parameters, and dynamically loading the corresponding front-end component modules according to the identifiers to render a labeling interface adapted to different data modalities.
[0180] Example 9, a mechanism for injecting multimodal pre-annotation results: When the annotation interface is initialized, the backend AI model inference process is automatically triggered, and the returned pre-annotation results with metadata are mapped to the front-end UI control to realize human-computer collaborative annotation.
[0181] Example 10, Unified Empty Shell Framework Structure: This framework decouples the data storage layer, the interface rendering layer, and the AI service layer.
[0182] Example 11: Low-code drag-and-drop layouts can be used instead of code configurations to achieve the same interface customization effect.
[0183] Example 12: In terms of AI model access, a locally deployed open-source small model can be used instead of a cloud-based large model. Although the accuracy may decrease, the pre-labeling function can still be achieved.
[0184] Figure 8 This is a block diagram of an electronic device structure based on a microkernel architecture multimodal dynamic annotation method provided by one or more embodiments of the present invention.
[0185] like Figure 8 As shown, the present invention provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0186] The memory stores a computer program, which, when executed by the processor, causes the processor to perform steps of a multimodal dynamic annotation method based on a microkernel architecture.
[0187] The present invention also provides a computer-readable storage medium storing a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform steps of a microkernel-based multimodal dynamic annotation method.
[0188] The communication bus mentioned in the above electronic devices can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not indicate that there is only one bus or one type of bus.
[0189] The electronic device comprises a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on the operating system. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory. The operating system can be any one or more computer operating systems that control the electronic device through processes, such as Linux, Unix, Android, iOS, or Windows. Furthermore, in this embodiment of the invention, the electronic device can be a smartphone, tablet computer, or other handheld device, or a desktop computer, portable computer, or other electronic device; there is no particular limitation in this embodiment.
[0190] In this embodiment of the invention, the executing entity for electronic device control can be an electronic device itself, or a functional module within an electronic device capable of calling and executing a program. The electronic device can obtain the firmware corresponding to the storage medium. This firmware is provided by the supplier, and different storage media may have the same or different firmware; no limitation is made here. After obtaining the firmware corresponding to the storage medium, the electronic device can write this firmware into the storage medium; specifically, it burns the firmware corresponding to the storage medium into the storage medium. The process of burning the firmware into the storage medium can be implemented using existing technology, and will not be elaborated upon in this embodiment of the invention.
[0191] Electronic devices can also obtain reset commands corresponding to the storage media. The reset commands corresponding to the storage media are provided by the supplier. The reset commands corresponding to different storage media can be the same or different, and no restrictions are imposed here.
[0192] At this time, the storage medium of the electronic device is a storage medium on which the corresponding firmware has been written. The electronic device can respond to the reset command corresponding to the storage medium on which the corresponding firmware has been written, thereby resetting the storage medium on which the corresponding firmware has been written according to the reset command. The process of resetting the storage medium according to the reset command can be implemented by existing technology and will not be described in detail in this embodiment of the invention.
[0193] For ease of description, the above apparatus is described by dividing it into various units and modules according to their functions. Of course, in implementing this invention, the functions of each unit and module can be implemented in one or more software and / or hardware.
[0194] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. It should also be understood that terms such as those defined in general dictionaries should be understood to have the meaning consistent with their meaning in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined.
[0195] For the sake of simplicity, the method embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0196] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present invention.
[0197] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A multimodal dynamic annotation method based on a microkernel architecture, characterized in that, include: S1. Receive task configuration parameters input from the administrator; The task configuration parameters include: the target data modality type and the corresponding functional component identifier; S2. Parse the task configuration parameters and extract the corresponding functional component identifiers; dynamically load the matching front-end atomic functional components based on the preset plugin standard protocol; wherein, the corresponding functional components are dynamically loaded according to the configured template ID; S3. Using a preset empty shell framework, the loaded functional components are laid out and rendered to generate a labeling interface that adapts to the target data modality type. S4. For the target data modality of the current annotation task, call the matching AI model service, obtain the AI pre-annotation results and inject them into the annotation interface; S5. Manually correct the pre-annotation results, receive the corrected annotation data, and complete data verification and persistent storage.
2. The multimodal dynamic annotation method based on microkernel architecture according to claim 1, characterized in that, include: A basic empty shell framework without business logic is used for the empty shell framework; Based on the aforementioned empty shell framework, it provides standard HTML container, event bus, layout management and basic data adaptation capabilities; All labeling business functions are implemented through callable functional components.
3. The multimodal dynamic annotation method based on a microkernel architecture according to claim 1, characterized in that, include: The preset plugin standard protocol defines a unified IAnnotationPlugin interface lifecycle specification for each functional component; The lifecycle specification includes: initialization (init), rendering (render), and component destruction (destroy) lifecycle methods; Each functional component is deployed and dynamically loaded independently based on the aforementioned lifecycle method.
4. The multimodal dynamic annotation method based on microkernel architecture according to claim 1, characterized in that, include: The functional components communicate with each other via an event bus within the shell framework, using a publish-subscribe pattern. In this way, component collaboration is achieved by subscribing to preset events, decoupling the direct coupling and calls between various functional components.
5. The multimodal dynamic annotation method based on a microkernel architecture according to claim 1, characterized in that, include: The target data modality types include any one or more combinations of audio modality, text modality, and image / video modality; Different functional components and AI models correspond to different modal configurations; When the target data modality is an audio modality, load the AudioPlayer audio playback component, the TimestampInput timestamp input component, and the TranscriptEditor rich text editing component; The ASR speech recognition model is invoked to generate pre-annotated text results with timestamps, and words with confidence scores below a preset confidence threshold are highlighted to remind the human correction end to review them. When the target data modality is a text modality, load the QueryDisplay text display component, the Tagselector tag selection component, and the EntityRelationGraph entity relationship graph display component; The NLP semantic understanding model is invoked to complete text intent recognition and automatic slot filling, and the recognized slots are highlighted with differentiated colors. When the target data modality is an image / video modality, load the CanvasRenderer rendering engine, ShapeDrawer graphics drawing toolbar, ClassSelector object category label selection component, and TimelineScrubber video timeline component; Call the CV vision detection model to generate target coordinates and category pre-labeled boxes, which can be used to support manual drag-and-drop correction of boundaries and adjustment of category labels; When the target data mode is a multimodal combination mode, the corresponding multiple functional components are loaded synchronously; Multi-component timeline synchronization is achieved through an event bus, enabling multi-modal fusion annotation.
6. The multimodal dynamic annotation method based on microkernel architecture according to claim 1, characterized in that, It also includes mechanisms for handling AI service anomalies and for circuit breaker fault tolerance: Monitor the service call status of AI models, and trigger service degradation strategies when service timeouts, recognition confidence scores fall below a preset threshold, network anomalies occur, or a preset number of consecutive call failures occur. Service degradation strategies include: activating AI-assisted functions via circuit breakers, switching to a purely manual annotation mode, and caching local data using IndexedDB. Data will be automatically synchronized once the network is restored.
7. The multimodal dynamic annotation method based on microkernel architecture according to claim 1, characterized in that, It also includes data consistency guarantee mechanisms: The front-end annotation view is updated in real time using the OptimisticUI optimistic update method; If backend data storage fails, the rollback view rollback method is triggered, restoring the view to its state before the annotation was edited and displaying an error message.
8. A multimodal dynamic annotation system based on a microkernel architecture, characterized in that, The method for implementing the microkernel-based multimodal dynamic annotation method according to any one of claims 1-7 includes: an empty shell framework layer, a functional plugin pool, a backend service layer, an AI model service layer, and a data persistence layer; The empty shell framework layer is used to provide basic capabilities such as basic layout rendering, event bus communication, task configuration parsing, user authentication, and data persistence. The functional plugin pool stores various atomic pluggable annotation functional components and supports dynamic loading on demand. The backend service layer is used to parse task configurations, route and schedule AI models, and process data verification and storage requests. The AI model service layer integrates an ASR speech recognition model, an NLP semantic understanding model, and a CV visual detection model to output pre-annotated results with metadata for each modality. The data persistence layer is used to store the annotation task configuration and the final annotation data.
9. An electronic device, characterized in that, include: The processor, communication interface, memory, and communication bus are connected, with the processor, communication interface, and memory communicating with each other via the communication bus. The memory stores a computer program that, when executed by a processor, causes the processor to perform the steps of the microkernel-based multimodal dynamic annotation method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, include: The device stores a computer program executable by an electronic device, which, when run on the electronic device, causes the electronic device to perform the steps of the microkernel-based multimodal dynamic annotation method as described in any one of claims 1 to 7.