Multimodal research data recording method, and system, terminal and storage medium
By using a multimodal scientific research data entry system and large language models and multimodal large models for structured processing, the system solves the problem of low efficiency in traditional experimental recording methods, realizes automatic identification and recording of multimodal data, and improves data accuracy and consistency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- WESTLAKE UNIV
- Filing Date
- 2025-05-21
- Publication Date
- 2026-07-23
AI Technical Summary
Traditional experimental recording methods are inefficient, complex to operate, and lack strong correlation and integration of multimodal data, which affects the comprehensiveness and accuracy of the recording.
The system employs a multimodal scientific research data entry system. It receives multimodal scientific research data through an interactive front-end module, performs structured processing using a large language model and a multimodal large model, generates structured operation instructions, automatically recognizes and records data, and supports voice interaction to correct errors.
It improves the efficiency of experimental recording, ensures the accuracy and consistency of data, reduces manual input time, and adapts to diverse scientific research needs.
Smart Images

Figure CN2025096189_23072026_PF_FP_ABST
Abstract
Description
Multi-modal scientific research data recording method, system, terminal and storage medium TECHNICAL FIELD
[0001] The present application relates to the technical field of data analysis, in particular to a multi-modal scientific research data recording method, system, terminal and storage medium. BACKGROUND
[0002] Although the traditional experimental recording method can import text data, voice data, photographic handwritten records and standardized data output by experimental instruments and equipment, etc., these data still need to be manually input or imported into the recording platform in a relatively passive manner, which is inefficient and inconvenient, and has the problems of low recording efficiency and complex operation. Especially in the case of busy experimental process or the need for rapid recording, the convenience of existing tools is difficult to meet the needs.
[0003] Moreover, although multiple modal data can be supported, the association and integration between different modal data are not close enough, which affects the comprehensiveness and accuracy of experimental recording. In addition, the traditional experimental recording is strictly written according to the format of the template, which is not flexible. SUMMARY
[0004] The present application is proposed in view of at least one of the above technical problems in the prior art. According to a first aspect of the present application, a multi-modal scientific research data recording method is provided, applied to a multi-modal scientific research data entry system, the multi-modal scientific research data entry system comprising a backend module; the method comprising:
[0005] Based on the current scientific research unit scheme, multi-modal scientific research data collected by an input device is received through an input interface; wherein the multi-modal scientific research data comprises at least one of the following data forms: text data, voice data, image data and structured data;
[0006] Based on the data field format specification, the backend module is called to structure the multi-modal scientific research data to convert the multi-modal scientific research data into structured operation instructions;
[0007] The structured operation instructions are converted into structured scientific research records according to the scientific research unit scheme.
[0008] The multi-modal scientific research data recording method of the present application embodiment can reduce manual input time, improve experimental recording efficiency, automatically identify and record multi-modal scientific research data, and correct errors through dialogue interaction, thereby ensuring data accuracy and consistency.
[0009] The above description is only a summary of the technical solutions of the present application. In order to make the technical means of the present application more clearly understood and implemented according to the content of the description, and in order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the following will describe the specific embodiments of the present application.
[0010] It should be understood that the foregoing general description and the following detailed description are only illustrative and explanatory, and are not limiting of the claimed application. BRIEF DESCRIPTION OF DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0012] Fig. 1 shows a schematic flow chart of a multi-modal scientific research data recording method according to an embodiment of the present application;
[0013] Fig. 2 shows a schematic flow chart of step S102 according to an embodiment of the present application;
[0014] Fig. 3 shows a schematic diagram of a multi-modal scientific research data recording method using an encoder to pre-process the multi-modal scientific research data according to an embodiment of the present application;
[0015] Fig. 4 shows a schematic flow chart of step S102 according to another embodiment of the present application;
[0016] Fig. 5 shows a schematic diagram of a multi-modal scientific research data recording method without using an encoder to pre-process the multi-modal scientific research data according to another embodiment of the present application;
[0017] Fig. 6 shows a schematic flow chart of using a first display area to display unstructured text data according to an embodiment of the present application;
[0018] Fig. 7 shows a schematic diagram of playing unstructured text data in a voice interaction mode according to an embodiment of the present application;
[0019] Fig. 8 shows a schematic flow chart of step S102 according to an embodiment of the present application;
[0020] Fig. 9 shows a schematic flow chart of step S801 according to an embodiment of the present application;
[0021] Fig. 10 shows a schematic diagram of converting text data into structured data field operation instructions based on a scientific research unit scheme prompt word according to an embodiment of the present application;
[0022] FIG. 11 shows a schematic diagram of converting picture data into structured data field operation instruction by picture encoder based on research unit scheme prompt word according to an embodiment of the present application;
[0023] FIG. 12 shows a schematic diagram of a display interface according to an embodiment of the present application;
[0024] FIG. 13 shows a schematic flow chart of step S103 according to another embodiment of the present application;
[0025] FIG. 14 shows a schematic flow chart of converting structured operation instruction into structured operation instruction mark code according to an embodiment of the present application;
[0026] FIG. 15 shows a schematic flow chart of checking data record based on data checking relationship and / or data constraint relationship between / among multiple data fields according to an embodiment of the present application;
[0027] FIG. 16 shows a schematic flow chart of permanently storing and releasing data record according to an embodiment of the present application;
[0028] FIG. 17 shows a schematic flow chart of managing content of research unit scheme data field input box according to an embodiment of the present application;
[0029] FIG. 18 shows a schematic flow chart of recording experiment in predetermined national language according to an embodiment of the present application;
[0030] FIG. 19 shows a schematic diagram of interaction between operator, AI interactive front end, data record end, AI back end and AI database according to an embodiment of the present application;
[0031] FIG. 20 shows a schematic block diagram of multi-model state research data entry system according to an embodiment of the present application;
[0032] FIG. 21 shows a schematic block diagram of terminal supporting multi-model state research data entry according to an embodiment of the present application. DETAILED DESCRIPTION
[0033] In order for those skilled in the art to better understand the technical solutions of the embodiments of the present application, the technical solutions of the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0034] Experimental records in the field of natural sciences are usually composed of data of different sources, different modalities and different representations (e.g., handwritten records, electronic spreadsheets, instrument photos, etc.). Meanwhile, in some special environments and scenarios, researchers need to be able to use voice interaction and free-form dialogue in an unlimited scenario to complete experimental records. However, data sharing and reuse of different modalities are difficult.
[0035] Based on at least one of the foregoing technical problems, the present application provides a multi-modal scientific research data recording method applied to a multi-modal scientific research data entry system, the multi-modal scientific research data entry system comprising an interactive front-end module and a back-end module: the interactive front-end module is mainly used for human-computer interaction; the back-end module is mainly used for structured processing and storage of multi-modal scientific research data. The method comprises: based on a current scientific research unit scheme, receiving multi-modal scientific research data collected by an input device through an input interface of the interactive front-end module; wherein the multi-modal scientific research data comprises at least one of the following data forms: text data, voice data, image data and structured data; based on the scientific research unit scheme and the data field format specification, calling the back-end module to perform structured processing on the multi-modal scientific research data, so as to convert the multi-modal scientific research data into structured operation instructions; and converting the structured operation instructions into structured scientific research records according to the scientific research unit scheme. The multi-modal scientific research data recording method of the present application embodiment can reduce manual input time, improve experimental record efficiency, automatically identify and record multi-modal scientific research data, and correct errors through dialogue interaction, thereby ensuring data accuracy and consistency.
[0036] The execution subject of the present application embodiment can be a multi-modal scientific research data entry system. The multi-modal scientific research data entry system can be an artificial intelligence (AI) system, which can take large language models (LLMs) or large multimodal models (LMMs) as the core, and the three are not distinguished in this paper. The AI system can realize structured processing of multi-modal scientific research data; it can not only automatically record experimental data, but also interact with the operator, correct numerical and logical errors, etc., thereby ensuring data accuracy and consistency.
[0037] The large language model herein refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. The large language model can handle various natural language tasks such as text classification, question answering, dialogue, etc. Currently, the large language model adopts a similar Transformer architecture and pre-training target as the small model, and the difference from the small model is to increase the model size, training data and computing resources. The large language model (LLM) of the embodiments of the present application can further include a plurality of encoders corresponding to a plurality of modalities, each encoder can convert data of the corresponding modality into text or its equivalent representation (such as token or vector embedding), so that the large language model has the ability to understand and process multi-modal information. The large language model can thus process multi-modal scientific research data and generate a series of executable structured operation instructions therefrom.
[0038] The multi-modal large model herein refers to training a large model in combination with text, image, audio, video and other multi-modal data, so as to realize the fusion and understanding of multi-modal information in a unified representation space; with this fusion capability, the multi-modal large model can not only process text-level semantics, but also can make more in-depth analysis and reasoning on data in non-text form such as image, audio, video, etc., and thus exhibit more comprehensive intelligent processing capability in complex tasks such as cross-modal retrieval, multi-modal question answering, image-text generation, etc. Under the paradigm of the multi-modal large model, the information contained in the multi-modal input can be directly processed by the multi-modal large model, and a series of executable structured operation instructions can be generated therefrom.
[0039] FIG. 1 shows a schematic flowchart of a multi-modal scientific research data recording method according to an embodiment of the present application; as shown in FIG. 1, the multi-modal scientific research data recording method 100 according to the embodiment of the present application can include the following steps S101, S102 and S103:
[0040] In step S101, based on the current Research Unit Protocol, the multi-modal scientific research data collected by the input device is received through the interactive front-end module input interface.
[0041] The multi-modal scientific research data includes at least one of the following data forms: text data, voice data, image data and structured data. The multi-modal scientific research data of the embodiments of the present application can provide more intuitive and detailed descriptions of research activities for the operator.
[0042] Text data in multimodal scientific research data can be text data input through an input interface by operators in the form of human-computer dialogue with a large language model. This text data can be text represented in natural language. Speech data can be a piece of speech input by operators through an input interface, or it can be in the form of audio or recordings. Image data can be images collected from experimental subjects, pictures taken from handwritten text, images of equipment, or video data. Structured data can be structured data generated by experimental instruments / equipment, or structured data exported from instruments / equipment. Structured data usually has a clearly defined data model and structure, can be stored in relational databases, and can be queried and manipulated using SQL (Structured Query Language). The characteristics of structured data are high organization; data is stored in rows and columns, has a fixed format and length, and each field has a predefined data type, such as integers, strings, dates, etc. The advantages of structured data include ease of retrieval, updating, and deletion operations, as well as data consistency and accuracy. Common types of structured data include numerical values, dates, times, telephone numbers, and addresses. Structured data can also include structured data generated by instrument application programming interfaces (APIs) / computational models, etc. In this embodiment, the operator can directly input structured data through the interactive front-end module.
[0043] The input interface can be a series of human-machine interfaces defined by the operator, through which multimodal scientific research data can be received.
[0044] The input device includes at least one of the following: keyboard, mouse, microphone, scanner, camera, handwriting tablet, drawing tablet, USB interface, network card, etc.
[0045] The research unit scheme here is essentially composed of components such as research protocols, models, and assigners. Within the research unit scheme protocol, operators can define various types of data fields and their corresponding data field format specifications (Data Field JSON Schema).
[0046] In step S102, based on the scientific research unit scheme and data field format specifications, the backend module is invoked to process the multimodal scientific research data, so as to convert the multimodal scientific research data into structured operation instructions.
[0047] Here, structured operation instructions refer to a series of single or multiple executable operation instructions.
[0048] In this embodiment, multimodal research data involved in the current research unit scheme can be sent to the backend module for processing. The large language model in this embodiment may include a speech encoder, an image encoder, a structured data encoder, and other encoders. Each modality of data corresponds to its own encoder. For example, speech data corresponds to a speech encoder, and image data corresponds to an image encoder.
[0049] It is worth noting that after parsing various types of data, the aforementioned encoders convert these data into text data according to the research unit scheme. The large language model then processes this text data to obtain structured operation instructions. Furthermore, since the encoders parse various data into text, no additional encoder is needed for parsing the text data in multimodal research data; instead, the text data can be directly sent to the large language model for processing.
[0050] In one embodiment of this application, as shown in FIG2, step S102 includes steps S201, S202 and S203:
[0051] In step S201, the multimodal scientific research data is identified and classified by the data modality recognition module to determine the type of the multimodal scientific research data;
[0052] In step S202, based on the data type of the multimodal scientific research data, the encoder and large language model corresponding to the data type of the multimodal scientific research data in the backend module are called to convert the multimodal scientific research data into the structured operation instructions;
[0053] In step S203, the back-end processing module converts the structured operation instructions into structured scientific research data.
[0054] The structured operation instructions in this application embodiment can consist of multiple directly executable operation instructions, and when multiple operation instructions are executed, they can generate structured data. The structured data will be stored in the input boxes of the scheme data fields of each scientific research unit, and finally form an experimental record.
[0055] For example, the structured operation instructions expressed in a certain programming language are as follows:
[0056] Since, under the research unit paradigm, the structured research data in a research record can actually be stored as a series of key-value pairs, the structured research data generated based on the above operation instructions is as follows:
[0057] In other words, the experimenter was changed to "Zhang San", the experimental temperature was changed to 25 degrees Celsius, and the experimental humidity was changed to 50%.
[0058] The core technical solution of this application is to convert the data in multimodal scientific research data such as text / voice / image input into structured operation instructions. In this way, the system can insert the relevant information of the structured scientific research data contained in these structured operation instructions into the corresponding data fields.
[0059] In one embodiment of this application, multimodal scientific research data can be structured according to its data type. Figure 3 illustrates an embodiment of this application where multimodal scientific research data is processed using an encoder corresponding to the data type. Examples of the processing procedures for different types of multimodal scientific research data are provided below.
[0060] In the first example, the multimodal research data is speech data. First, a speech encoder corresponding to the speech data is invoked to convert the speech data into text. Then, the text is output. For example, referring to Figure 12, when the operator says "Search PCR," the speech encoder converts the speech data into the text form "Search PCR" and displays it in the first display area.
[0061] In the second example, the multimodal scientific research data is image data. First, an image encoder corresponding to the image data is invoked, so that the image encoder converts the image data into text based on an annotated industry knowledge dataset using a text recognition algorithm and / or an optical character recognition algorithm. Then, the text is output.
[0062] Continuing with Figure 3, for example, image data can include photos of handwritten text and photos of devices. For handwritten text images, a text recognition algorithm (Texas Extract algorithm) can be used to extract the text from the handwritten text image, recognizing it as text. For experimental objects, a text recognition algorithm (Texas Extract algorithm) and an optical character recognition (OCR) algorithm can be used to analyze the device image, extracting the text and target objects, and then generating text. For device images, an optical character recognition (OCR) algorithm can be used to analyze the device image, obtaining the device's status, and then generating text. For example, if image analysis reveals that a device's lid is not properly closed, this can be recorded in the experimental record as "The lid of a certain device is not properly closed." This method not only directly converts image data into text on the terminal but also provides prompts for non-standard experimental operations, thereby improving the accuracy of the experiment.
[0063] Referring to Figure 12, for example, before the formal experiment, the operator needs to record the experimental environment. For instance, a photo of a certain experimental device is taken and then input into the AI system. The AI system's backend analyzes the photo and finds that the lid of a certain experimental device is not properly closed. The AI system will record this information and display it as a text annotation in the second display area 1202, namely, "The lid of a certain device is not properly closed".
[0064] It is worth noting that video data is composed of multiple frames of images, and video data belongs to image data. Therefore, the processing of video data is the same as that of image data.
[0065] Image data is an important type of unstructured data source. In scientific experiments, it is often necessary to photograph the experimental process, and operators sometimes need to handwrite experimental records. This application embodiment will use a large language model to recognize image-based scientific research data and store it in the Research Unit Scheme Data Field Input Box (RU Data Field Input Box). The processing of image data mainly falls into two categories: First, during the initialization phase, image information is used for initialization to identify and record the scientific research data contained within the image information; second, consistent with the scenario mentioned above regarding experimental records of text data, image data can be input during dialogue.
[0066] In the fourth example, the multimodal scientific research data is structured data. This application also supports experimental records in structured marked documents such as Word, JSON, and Markdown, which can be identified and entered using large language models and structured parsing tools.
[0067] In this embodiment, the operator can manually fill in structured data in the research unit plan data field input box according to the current research unit plan. For example, the operator can manually fill in experimental data such as experimenters, experiment name, experiment date, experiment temperature, experiment objects, and changes in the experimental objects in the research unit plan data field input box. It is worth noting that when manually filling in structured data, the data can be directly entered into the research unit plan data field input box in the second display area so that the second display area displays the structured data. The data type and value of the structured data should follow the rules defined in the research unit plan data field input box.
[0068] Regarding experimental records output by instruments, equipment, and computational models, since instruments, equipment, and computational models widely used in modern scientific experiments can automatically generate a large amount of data, such as experimental measurements, instrument status, and statistical results, embodiments of this application can enable large language models to interface with the APIs of such instruments, equipment, and models to obtain input and output information and record experiments.
[0069] Furthermore, since the structured operation instructions generated by the large language model are text-based, when acquiring data such as voice data, image data, and structured data, it is necessary to convert this data into text (e.g., voice, image, and video data encoded in Base64 string format; or upload voice, image, and video data to a local database / cloud database / LAN / Internet to obtain the corresponding text-based ID / URL / URI). This text is then stored in the research unit scheme data field input box and displayed on the display interface. If the structured operation instructions generated by the large language model contain voice data, this voice data can be stored in the recording slot and played upon receiving a playback command. If the structured operation instructions generated by the large language model contain image data, this image data can be stored in the recording slot and displayed on the display interface or preview window upon receiving a display command. If the structured operation instructions generated by the large language model contain video data, this video data can be stored in the recording slot and played through a video player upon receiving a playback command. If the structured operation instructions generated by the big language model contain data in the form of structured documents, these structured document data can be stored in the record slots, and automatically parsed and presented according to the preset format requirements when parsing or display instructions are received.
[0070] The embodiments of this application can realize intelligent experimental recording on a general scalable platform, that is, automatically identify and record data such as text data, voice data (including voice commands), handwritten text images, and standardized data output by experimental instruments and equipment.
[0071] In another embodiment of this application, as shown in FIG4, step S102, based on the data field format specification, calls the backend module (including the multimodal large model) to perform structured processing on the multimodal scientific research data, so as to convert the multimodal scientific research data into structured operation instructions (single / multiple operation instructions), including steps S401 and S402:
[0072] In step S401, the backend module is invoked, and its built-in algorithm is used to directly convert the multimodal scientific research data into structured operation instructions.
[0073] In step S402, the structured scientific research data contained in the structured operation instruction is stored in the scientific research unit scheme data field input box and displayed in the second display area.
[0074] As shown in Figure 5, after receiving multimodal research data (text data, voice data, image data, instrument API / computational model data, and other data), the input device directly sends the multimodal research data to the multimodal large model, allowing the multimodal large model to process it using its built-in algorithms. For example, the multimodal large model can use deep learning algorithms to analyze and understand the multimodal research data, generate multiple structured operation instructions, and automatically execute these instructions to generate corresponding structured text data. This structured text data will be stored in the research unit scheme data field input box. Additionally, after analyzing and understanding the multimodal research data, the multimodal large model can output unstructured text data and display the structured text data on the display interface (e.g., the first display area) to allow users to intuitively understand the experimental record content.
[0075] In one example, referring to Figure 12, the second display area 1202 in Figure 12 displays structured text data obtained after executing structured operation instructions. This structured text data will be presented in the final experimental record. Users can directly enter the corresponding content, i.e., structured text data, into the research unit scheme data field input boxes corresponding to data fields such as "Experiment Personnel," "Experiment Date," "Experiment Purpose," "Temperature," and "Humidity."
[0076] In another example, continuing with Figure 12, the first display area 1201 in Figure 12 displays unstructured text data. This data is obtained by the AI system through analysis and understanding of multimodal data (such as voice, images, and text). This unstructured text data needs to be processed again to be presented as structured text data as shown in the second display area 1202. For example, if an operator asks the question "Please change the temperature to 45 degrees" via voice, the AI system, after analyzing and understanding the question and performing the operation, will reply via voice "The temperature has been changed to 45 degrees."
[0077] In other words, in this example, the AI system will convert multimodal data into unstructured text data (such as answers generated by the AI based on user questions) and present it in the first display area 1201. It is worth noting that users can input questions in the form of text data. For example, if a user enters the question "Search PCR" in the question input box in the first display area 1201, the AI system, after searching, will not find PCR and will display the text message "Unable to find PCR".
[0078] It is worth noting that the current research unit scheme not only performs intuitive language conversion for the input multimodal research data, but also uses algorithms such as deep learning to understand the input data. Problems encountered during the actual process can be recorded and displayed in the final experimental record. For example, if parsing a photo (or video) of a device reveals that the lid is not properly closed, the research unit scheme data field input box will record "The lid of the device is not properly closed." This application embodiment can parse photos, obtain the state of objects in the photos, and record the state of the objects in the research unit scheme data field input box to facilitate timely adjustments by operators.
[0079] The AI system in this application embodiment can better fuse, analyze, and understand multimodal data. For example, for image and audio data acquired during an experiment using a camera and microphone, image encoders and speech encoders can be used to convert the data into text, which can then be analyzed and understood using a large language model, or a multimodal large model can be used directly for analysis and understanding to generate corresponding experimental records. Furthermore, because the understanding provided by the large language model / multimodal large model is highly accurate and flexible, it can answer experiment-related questions, assisting operators in recording experimental data.
[0080] As shown in Figure 6, the method further includes steps S601 and S602:
[0081] In step S601, unstructured text data is sent to the first display area so that the first display area displays the unstructured text data in the form of text blocks; and / or
[0082] In step S602, an unstructured text instruction input by the operator in the first display area is received, and a response is made to the unstructured text instruction.
[0083] In one embodiment of this application, the multimodal scientific research data entry system further includes an interactive front-end module. As shown in Figure 12, the display interface 1200 of the interactive front-end module includes a first display area and a second display area. The first display area is used to display human-computer question-and-answer content in text form, or to display voice data (which has been parsed into text data by a voice encoder), image data (which has been parsed into text data by an image encoder), or structured data (which has been parsed into text data by a structure encoder). For a detailed description of the first and second display areas, please refer to the following text.
[0084] For example, Natural Language Processing (NLP) technology can be used to perform speech-to-text conversion on the voice interaction data to generate text corresponding to the speech; the voice interaction data may include voice commands, requests for the AI system to record scientific research data, and / or answers to questions related to the scientific research unit plan and recording key information during the experimental process.
[0085] In some embodiments, the content displayed in the first display area can be played aloud simultaneously with the display, allowing operators to access the displayed content and AI system feedback without having to physically approach the display interface. Furthermore, users can directly provide voice commands to the broadcast, enabling the AI system to process and provide further feedback based on the user's additional voice commands. As shown in Figure 7, the method further includes steps S701 and S702:
[0086] In step S701, the unstructured text data is played using voice interaction; and / or
[0087] In step S702, the operator's voice command is received and responded to.
[0088] In human-computer multi-turn question-and-answer (voice interaction), the large language model can use natural language processing technology to perform operations such as speech-to-text conversion and text summarization on the voice interaction data. Operators can input experimental records by voice, and the large language model will convert the speech into text and store it. Through the NLP-based dialogue system, it can interact with the operator and answer experimental-related questions, and help record key information in the experimental process.
[0089] In some embodiments, the operator can also issue operational instructions to the large language model. After executing these instructions, the AI system can provide feedback to the operator on the execution status. For example, when starting an experiment, the operator says "Please initialize," and the conversation module in the AI system (e.g., a chatbot) obtains the initialization status of each experimental instrument or device based on the generated text summary "initialize." When all experimental instruments or devices have completed the initialization operation, the system responds with a voice message, "Initialization complete." As another example, if the operator says "Experiment X, Zhang San," the AI system can understand "Zhang San" as either the experimental operator or the experimental recorder based on natural language processing algorithms, and then record "Operator: Zhang San" or "Experimental Recorder: Zhang San" in the experimental record.
[0090] In addition, based on the operator's custom rules, the AI system can be configured to maintain a human-computer question-and-answer process throughout the entire experiment, so as to record data generated by multiple experimental steps in a single experiment or multiple scientific research data from multiple experiments.
[0091] Furthermore, the AI system can not only answer operators' questions, but also search based on keywords related to the questions. In this embodiment, operators can issue instructions to the AI system in a multi-turn dialogue to perform operations such as modifying, correcting, deleting, storing, and updating data based on the experimental process data displayed on the interface.
[0092] For example, when an operator says "Search PCR for me," using "PCR" as the field name, and no "PCR" is found, the system can provide text / voice feedback such as "PCR not found, please try again." The system can be set to allow three searches before responding, or to stop searching if no results are found after three attempts. This embodiment of the application can perform semantic analysis on the voice data, execute instructions based on their meaning, and then convert the execution results into text, providing feedback to the operator in the form of voice data.
[0093] Furthermore, operators can engage in multiple rounds of question-and-answer sessions with the AI system throughout the experiment, and the question-and-answer process and results can be stored and displayed in the research unit scheme data field input box. Addressing the issue that operators may not be able to free their hands to record research data in certain specific scenarios, this application embodiment can also record the experimental process based on voice interaction. The AI system converts the operator's voice input and model output into the aforementioned scenario of recording text data in the experimental record, thus completing the experimental record based on voice dialogue commands.
[0094] For example, a user can give the system a voice command: "Please start recording. Record the experimenter as Zhang San." The system will reply: "Experimenter recorded as Zhang San," and announce it verbally. After hearing the voice message, the user can continue to give a voice command: "Please record the experiment date as January 1, 2024." The system will then reply: "Experiment date recorded as January 1, 2024." Thus, users can record scientific data directly through the voice control system without needing to approach the recording terminal or even view the interface. This method effectively addresses situations where users cannot use their hands for keyboard operation, such as when conducting a chemical experiment inside a glove box and needing to operate continuously inside the glove box without being able to use a keyboard to input data.
[0095] Therefore, the embodiments of this application can adapt to the diverse needs of scientific research data and effectively improve the ease of operation and accuracy of researchers in the recording process.
[0096] In step S102, the unstructured text instructions can also be converted into structured operation instructions according to the research unit scheme.
[0097] As shown in Figure 8, step S102, which converts unstructured text instructions into structured operation instructions according to the research unit scheme, includes step S801:
[0098] In step S801, based on the prompt words of the current research unit scheme, the unstructured text instructions are analyzed and understood using a preset algorithm to convert the unstructured text data into corresponding structured operation instructions.
[0099] For example, a user issues the following unstructured text instruction:
[0100] "Please set the temperature to 25℃."
[0101] Using S801, structured operation instructions can be converted into the following form:
[0102] In one embodiment of this application, the encoder converts the multimodal research data into text based on the current research unit scheme prompt.
[0103] The prompts for the research unit plan here can be automatically generated based on the current research unit plan, or preset by the operator or other personnel according to the current research unit plan. For example, the prompt "Temperature range is 20℃-50℃" can be displayed below the experimental temperature field. When the obtained temperature is not within this temperature range, a temperature error prompt will be sent.
[0104] As shown in Figure 9, step S801, based on the prompt words of the current research unit scheme and using a preset algorithm to analyze and understand the unstructured text data to convert the unstructured text data into corresponding structured operation instructions, further includes step S901:
[0105] In step S901, an embedded tool specific to the research unit scheme prompt word architecture is used to vectorize each knowledge point in each text block and store it in the research unit scheme data field input box in the form of key-value pairs for subsequent fast matching indexing.
[0106] Wherein, the research unit scheme prompt words are pre-set in the backend module; and / or, the research unit scheme prompt words are automatically created by the backend module based on the knowledge base corresponding to the current research unit scheme.
[0107] In one embodiment of this application, Figures 10 and 11 are used to illustrate how to convert unstructured text data into corresponding structured operation instructions based on research unit scheme prompts. Figure 10 shows a schematic diagram of converting text data into structured operation instructions based on research unit scheme prompts according to an embodiment of this application; Figure 11 shows a schematic diagram of converting image data into structured operation instructions based on research unit scheme prompts according to an embodiment of this application. The research unit scheme prompts are created by the AI system based on the current research unit scheme. The AI system can also construct its own research unit scheme background prompts based on the background knowledge of the experiment. Background knowledge may include basic information defined by the operator, working environment, experimental tasks, experimental objectives, etc. For example, before conducting a medical experiment to treat a lung disease, basic information such as room temperature and whether the work is carried out in a sterile environment can be input into the AI system, from which the current research unit scheme background prompts can be extracted and added to the existing research unit scheme prompts.
[0108] Figure 12 shows a schematic diagram of the display interface 1200 of the multimodal scientific research data recording method according to an embodiment of this application. The display interface 1200 implemented in this application includes a first display area 1201 and a second display area 1202. The first display area 1201 is located on the left and is used to display the content of the human-computer interaction question and answer; the second display area 1202 is located on the right and is used to display the content temporarily stored in the scientific research unit scheme data field input box. In addition, the first display area 1201 and the second display area 1202 are respectively provided with scroll bars, so that users can manually / automatically scroll the scroll bars to display more content when the human-computer interaction question and answer content and the scientific research unit scheme data field input box and its content are large.
[0109] As shown in Figure 13, step S103, which converts the structured operation instructions into structured research records according to the research unit scheme, also includes step S1301:
[0110] In step S1301, the structured data is stored in the scientific research unit scheme data field input box and displayed in the second display area.
[0111] Referring to Figure 12, since different data types (such as strings, integers, floating-point numbers, booleans, dates, enumeration values, etc.) are defined for different fields in the data field format specification, different research unit scheme data field input boxes can be generated for the data fields in the second display area. Furthermore, corresponding interactive controls can be generated for the research unit scheme data field input boxes based on the data type specifically annotated for each field in the data field format specification. For example, data type-based annotations can be provided for the research unit scheme data field input boxes based on the data type rules of the data field format specification. This embodiment of the application, through this display method, helps operators concentrate their time and energy on defining and developing the substantive content of the research scheme, such as research agreements and data fields, without needing to worry about issues such as experimental record interfaces, research data storage structures and methods. This enables scientists to efficiently design high-quality research schemes that meet actual research needs in a user-friendly manner during daily research activities and use them for research data recording.
[0112] All research data are stored in the database in the form of fields. It is worth noting that, in this embodiment, because the AI system can directly obtain research data output from various experimental instruments or equipment, all data in each field on the right side of the display interface can be directly displayed without manual input, thus improving efficiency compared to traditional manual input methods. Furthermore, this embodiment displays all data in a visual format, allowing operators to easily access the research data at any time. For situations with a large amount of data on the right, the research data can be automatically organized, and an automatic scrolling interface can be set, further improving experimental recording efficiency.
[0113] As shown in Figure 14, the method further includes steps S1401 and S1402:
[0114] In step S1401, the backend module converts the structured operation instructions into code marked with the structured operation instructions.
[0115] In step S1402, the code marked with the structured operation instructions is executed to form a structured scientific research record.
[0116] The code marked by the structured operation instruction can be represented as multiple operation instructions expressed in a certain machine language that can be automatically executed by the terminal.
[0117] In one embodiment of this application, the research unit scheme data field input box is further used to store data corresponding to multiple data fields. Here, the data corresponding to multiple data fields refers to data extracted from the corresponding data fields in multimodal research data such as text / voice / images input by the user.
[0118] For example, if a user sets the temperature to 45 degrees Celsius and the humidity to 50%, then the corresponding processed structured scientific data would be:
[0119] The 25.0 and 50.0 here are the corresponding data in the data field.
[0120] As shown in Figure 15, the method further includes step S1501:
[0121] In step S1501, based on the data verification relationship and / or constraint relationship between the multiple data fields, the data corresponding to the multiple data fields is verified, and if there is a logical error in the data corresponding to the multiple fields, a prompt indicating that there is a logical error is sent.
[0122] Specifically, the type constraint includes constraining the corresponding data field to use predefined multimodal research data during data entry. This predefined multimodal research data includes one or more of text data, image data, video data, audio data, and text data. In some embodiments, the operator can further define the data type (e.g., various numeric types, time types, etc.) of the data field previously defined in the research agreement. For example, if the `solvent_volume` data field is defined as a floating-point number, then when the operator enters non-floating-point data (e.g., a string of letters) into this field, the system will indicate a type error.
[0123] The numerical validation relationship includes constraining the corresponding data fields to follow a specified pattern and / or not exceed a preset value range during data entry. Operators can further define validation rules for the data fields defined in the research agreement. For example, they can add validation rules to the `solvent_volume` data field and its type constraint (floating-point number) to ensure that the floating-point number entered into this data field must be greater than zero. In this case, if the operator enters a negative floating-point number, the system will display a numerical validation error. In another example, following a specified pattern can be, for example, using a RegExp regular expression (RE) to constrain the composition rules and patterns of strings in the field. For example, constraining a field related to email addresses to record values that contain exactly one "@" symbol, etc. Other patterns and value range constraints can also be set in other examples, which are not listed here.
[0124] The combined verification relationship includes constraining the types and / or values of each data field to meet predetermined constraints. For example, it can be constrained that in an experiment, if the required temperature is greater than 40 degrees Celsius, the required humidity must be less than 20%. Thus, temperature-related values and humidity-related values form a combined verification relationship, which are interdependent.
[0125] This enables the rapid identification of abnormal research data. If any validation fails, the reason for the failure will be displayed, allowing operators to correct the values of invalid data fields one by one according to the error messages until a valid record is obtained. Through this validation process, the platform ensures that the data entered by operators meets the requirements.
[0126] In other embodiments, if an operator defines multiple data fields with dependencies and assignment relationships in a research agreement, these relationships can be further customized using assigners. Specifically, when an operator designs a research unit scheme for a corresponding discipline based on a research node design environment, it can also include assigners for data fields. These assigners are used by the operator to assign values to data fields based on the data field dependency graph and assignment rules. In some embodiments, the data field dependency graph is a single-level or multi-level directed acyclic graph, the assignment relationships between the defined data fields are single or multiple dependencies, and each data field is assigned a value by at most one assigner. An assigner can have one or more upstream data fields as dependencies, and a data field can actually be a dependency of one or more data fields, but for a specific data field, the method of determining its field value should be unique. As an example, if two data fields, solvent_volume and solvent_volume_2, are defined in a research protocol, and the operator uses an assigner to ensure that solvent_volume_2 is always twice the value of solvent_volume, then whenever a new value is entered for solvent_volume, the system will automatically set solvent_volume_2 to twice the solvent volume value. For example, if the value of solvent_volume is 5, the system will automatically assign the value 10 to solvent_volume_2. This can greatly improve the efficiency and accuracy of research data recording.
[0127] By defining a model for the research unit scheme, the input of operators can be dynamically verified to ensure the accuracy of data entry. By defining the assigner, multi-level and multi-dependent field dependencies can be automatically calculated based on the value of a certain input field, thereby ensuring that even if there are complex dependencies between multiple data fields, they can be efficiently entered and run correctly. This can significantly promote the electronic management of laboratory research data, including data from outsourced experiments and orders, as well as the efficient retrieval of research schemes, research data, and other related content.
[0128] For example, referring to Figure 12, the temperature and humidity data obtained by the Large Language Model (LLM) are 45 degrees Celsius and 70%, respectively. However, the humidity cannot be 70% when the temperature is 45 degrees Celsius. In this case, a prompt can be issued on the display interface, for example, by highlighting the humidity data to attract the operator's attention; an audible prompt can be issued using a buzzer; or the humidity data can be displayed as text below the humidity field to prompt the operator. This embodiment of the application can be configured to immediately verify each piece of data obtained, and after all data has been obtained, verification can be performed again based on the combined verification relationships between related data to ensure the accuracy of the experiment and experimental records.
[0129] This application embodiment transforms experimental records from different modalities into a data format (e.g., text) that can be recognized by a Large Language Model (LLM) through encoders corresponding to those modalities. Subsequently, the powerful dialogue and task understanding capabilities of current Large Language Models (LLMs) are utilized to update and record experimental records from different modalities. Furthermore, during dialogue with the large model, interactive input such as text and speech can be used to correct research data with recognition errors in other modalities. While completing the recognition and recording of research data, high-quality data is collected for training and fine-tuning the multimodal research large model.
[0130] In one embodiment of this application, as shown in FIG16, the method further includes steps S1601 and S1602:
[0131] In step S1601, based on the storage instruction input by the operator, after storing the structured research record in the database according to the research unit scheme, the data corresponding to the data field stored in the data field input box of the research unit scheme is released; or,
[0132] In step S1602, based on the non-storage instruction input by the operator, the data corresponding to the data field stored in the data field input box of the scientific research unit scheme is directly released.
[0133] In this embodiment, the data in the research unit scheme data field input box is not permanently stored, but rather temporarily stored as various data during the research / experiment process. Once the final structured research record is generated, the data in the research unit scheme data field input box will be cleared or migrated to other storage devices. For example, if the multimodal research data entry system is installed on a local terminal, the data stored in the research unit scheme data field input box is correspondingly stored in the local terminal's cache. When it needs to be stored as the final structured research record, the operator clicks the "Submit" button on the display interface, and the current structured research record will be stored in the local database. As another example, if the multimodal research data entry system is installed on a cloud server, the data stored in the research unit scheme data field input box is correspondingly stored in the cloud server. When it needs to be stored as the final structured research record, the operator clicks the "Submit" button on the display interface, and the current structured research record will be stored in the cloud database.
[0134] It is worth noting that when the structured research records are stored in the database, a structured storage scheme will be automatically generated for each research unit scheme based on the data field format specifications. This facilitates the submission and storage of structured research records based on the research unit scheme by the operators, and each data field should conform to the constraints and data validation relationships of the model. In the embodiments of this application, regardless of whether the data fields defined by the operators are simple text or more complex multimodal data, the large language model can automatically generate the corresponding data structure for them and ensure the consistency and integrity of the data during the input and storage process. Since the automatically generated data structure is based on standardized protocols and definitions, research data can be easily shared globally. This automation of structured storage not only simplifies the work of researchers but also improves the reproducibility of research data and the ability to collaborate across laboratories. Furthermore, with the automatically generated data structure, researchers can focus on experiments and data recording without worrying about underlying data management issues. This feature greatly reduces the data management burden on researchers and improves the efficiency of research work.
[0135] In one embodiment of this application, as shown in FIG17, the method further includes step S1701:
[0136] In step S1701, based on the management instructions input by the operator, the content stored in the input box of the research unit scheme data field is managed.
[0137] The management operations include at least one of the following: deleting data, storing data, and modifying data.
[0138] For example, operators can manually delete or modify experimental data in the data input field of the research unit scheme.
[0139] For example, when an operator discovers a problem with the temperature data in a set of scientific research data, they might say, "Correct the temperature in the third set to 45℃." The AI system receives the voice data and can directly process it using its built-in algorithm, converting the voice data into structured operation instructions.
[0140] Then, based on the above structured operation instructions, the AI system will fill the structured scientific research data (corresponding to the scientific research unit plan data field "temperature" with a value of "45℃") into the corresponding scientific research unit plan data field input box; or, the large language model will convert the speech data into unstructured text data through the speech encoder, and then convert the unstructured text data into structured text data, that is, the scientific research unit plan data field "temperature" with a value of "45℃" will be filled into the corresponding scientific research unit plan data field input box.
[0141] For example, if an error occurs in the temperature input during an experiment and needs to be corrected, the operator can say "Change the third group temperature to 45℃". After the AI system changes the value in the scientific research unit scheme data field input box of the temperature field to 45℃, the AI system can also provide feedback, such as returning the text "The third group temperature has been corrected to 45℃" in natural language, or further, broadcasting the text via voice, so that the operator can obtain feedback from the AI system through voice without having to directly look at the first display area.
[0142] In one embodiment of this application, as shown in FIG18, the method further includes steps S1801 and S1802:
[0143] In step S1801, an instruction is obtained to record the experimental records in a predetermined national language;
[0144] In step S1802, the structured operation instructions are translated into the predetermined national language and recorded in the experimental record.
[0145] For example, if the operator's native language is Chinese and the desired country language is English, the user can interact with the AI system using their preferred language. For instance, they could issue an instruction in Chinese: "Please record the weather as sunny" (this research unit solution contains a data field with the ID "weather"). The AI system can then automatically understand the user's instruction and generate the following structured instruction:
[0146] Since the language of the country预定 by the operator is English, the AI system can further translate the information in the above operation instructions into English (translating "sunny day" to "Sunny"), and fill in "Sunny" in the data field input box corresponding to "weather".
[0147] The embodiments of this application can support the input of multi-modal scientific research data during the Q&A process, and carry out the Q&A process based on non-text modal data such as pictures, videos or voices, expanding the usage scenarios of the Q&A process. Moreover, by expressing different forms of information through different modal data, the expression of information becomes more abundant, thus being able to provide more abundant information to the large model, helping to improve the accuracy of the large model's understanding of the input content, and thus generating more accurate reply content, enhancing the accuracy of the reply content. During the Q&A process, when the conversation content of the embodiments of this application includes non-text content, through multi-modal intent recognition of the conversation content and non-text content, the multi-modal intent recognition result indicates the content relevance between the conversation content and the non-text content, so that a suitable target reply model can be selected according to the content relevance, thereby being able to improve the accuracy of the selected reply model, and further enhancing the accuracy of the reply content.
[0148] The present invention will be introduced in detail again below with the AI model as an example in conjunction with FIG. 19.
[0149] As shown in FIG. 19, it is a schematic diagram of the interaction between an operator, an AI interaction front end, a data recording end, an AI back end, and an AI database.
[0150] It can be seen from FIG. 19 that the data of the embodiments of this application mainly interacts between an operator, an AI interaction front end, a data recording end, an AI back end, and an AI database. The input method can be divided into three main stages: the first stage is the data processing stage, the second stage is the output stage, and the third stage is the stage of saving the interaction information between the operator / AI.
[0151] In the first stage, the input multi-modal scientific research data is sent to the AI back end for processing to generate structured operation instructions, that is, multiple operation instructions, and multiple operation instructions are executed to obtain data records.
[0152] The steps are as follows:
[0153] Step 1, input relevant information of the data fields (DFs) to be recorded in the AI interaction front end in the form of dialogue / voice / picture, etc.
[0154] Step 2, the AI interaction front end sends the operator's instructions to the AI back end.
[0155] Step 3: Based on the scientific research unit scheme and data field format specifications (data field JSON Schema), AI processes the information input by the operator at the front end into a structured operation instruction in the form of a list on the data record end.
[0156] Step 4: The AI backend sends the operation instructions to the data recording terminal.
[0157] Step 5a: The data recording end processes each operation instruction one by one and generates Acknowledge information corresponding to each operation instruction.
[0158] Step 5b: For each operation command successfully processed by the data recording terminal, the data field input box filled in by the operation command and the value entered are displayed in real time on the data recording interface.
[0159] In the second stage, the AI can also generate a second AI output. The steps are as follows:
[0160] Step 6a: The data recording end transmits the list-style Acknowledge information to the AI backend.
[0161] Step 6b: The AI backend generates the AI second output, i.e., the response to the operator, based on a. the operator's original input, b. the AI's first output (list of operation instructions), and c. the data recording terminal response (list of acknowledgement information).
[0162] In step 6c, the AI backend sends the second AI output to the AI interaction frontend.
[0163] In step 6d, the AI interactive front end displays the information from the previous step to the operator (usually in the form of a dialogue response).
[0164] In the third stage, data records can be saved to a database, including the content of multi-turn human-machine question-and-answer sessions. This database can be a local database or a cloud-based database. The steps are as follows:
[0165] Step 7: The AI backend stores the following information into the AI database: a. the operator's original input, b. the AI's first output (list of operation instructions), c. the data recording terminal response (list of acknowledgement information), and d. the AI's second output.
[0166] In this embodiment, handwritten text photos and device photos (or videos in other embodiments) are first acquired, which can be manually input by the operator into the execution subject (e.g., a computer). Based on the research unit scheme prompts corresponding to the specific research unit scheme, the image encoder converts the handwritten text photos and device photos into corresponding descriptive text. A dialogue robot based on an AI system (the dialogue robot can be part of the AI system) outputs research data. This research data is stored in a temporary research unit scheme data field input box for temporary recording and modification of experimental results. After the operator confirms the results, the experimental record results of this experiment will be stored in the system's data terminal.
[0167] The data field input box of the research unit scheme allows for flexible control over the output format of experimental records. It also enables the model to modify and manually correct erroneous results. These manual correction operations will be collected as human feedback data, providing a data basis for the training and fine-tuning of the unified joint representation multimodal model.
[0168] It is worth noting that the processes from inputting handwritten text photos and device photos into the image encoder (which belongs to the large language model) to the data storage in the recording slot by the large language model-based dialogue robot (also belonging to the large language model (LLM)) can all be completed by the large language model (the various functions mentioned in this application can also be implemented by using a multimodal large model). This can help operators record scientific research data more quickly and reduce the workload of manual input.
[0169] The multimodal scientific research data recording method of this application can be applied to technical fields such as medical records, legal document management, and educational data analysis that require AI system recording and intelligent management. As an example, it can be directly applied to electronic medical record scenarios. Similar to the scientific research described in other embodiments of this application, in medical / hospital / clinical scenarios, various institutions / departments have a need for medical data recording (such as electronic medical records). However, the content, type, and specifications of the data to be recorded vary across different departments, diseases, and medical scenarios. It is conceivable that if each department records data in its own customized way, it will not only be inefficient but also hinder sharing and experience accumulation. With the help of the scientific research activity management and application platform in this application, operators from different professional fields can customize the protocols, models, and assigners of medical records according to the scientific research unit syntax. This allows for tailored medical recording methods for specific needs, and the customized scientific research unit scheme can be widely shared across different departments. This enables a single design for hospital-wide / cross-hospital application, achieving unification and standardization for specific medical records. Furthermore, in clinical medical data recording, doctor-patient consultations are frequently encountered where doctors may not have both hands free to operate computers or other devices, or to manually record clinical medical data. The multimodal data (e.g., voice-based) recording solution proposed in this application provides a convenient solution to this recording need. In summary, this application has significant application potential and substantial economic benefits across multiple industries.
[0170] The multimodal scientific research data recording method of this application converts the multimodal scientific research data of the current scientific research unit scheme into structured operation instructions, and then combines at least one round of question-and-answer process with the operator to convert the structured operation instructions into structured scientific research records according to the scientific research unit scheme. This can reduce manual input time, improve experimental recording efficiency, and utilize an AI system to automatically identify and record multimodal scientific research data, and correct errors through dialogue interaction, ensuring data accuracy and consistency.
[0171] Figure 20 is a schematic block diagram of a multimodal scientific research data entry system 2000 according to an embodiment of this application. As shown in Figure 20, the multimodal scientific research data entry system 2000 according to this application includes an interactive front-end module 2001 and a back-end module 2002.
[0172] The backend module 2002 is used to receive multimodal scientific research data collected by at least one input device through an input interface, based on the current scientific research unit scheme.
[0173] The multimodal scientific research data includes at least one of the following data formats: text data, voice data, image data, and structured data.
[0174] The backend module 2002 is also used to perform structured processing on the multimodal scientific research data based on the data field format specifications, so as to convert the multimodal scientific research data into structured operation instructions; and to convert the data into structured scientific research records according to the scientific research unit scheme.
[0175] Among them, structured operation instructions can consist of multiple operation instructions, and when multiple operation instructions are executed, they can generate structured data.
[0176] The interactive front-end module 2001 includes a first display area and a second display area. The first display area is used to display and / or play unstructured data, and the second display area is used to display structured data.
[0177] In this embodiment, no manual filling is required; instead, the experimental records are automatically filled in by the AI system, which improves efficiency compared to the traditional manual filling method.
[0178] The terminal supporting multimodal scientific data input according to this application will be described below with reference to FIG21, wherein FIG21 shows a schematic block diagram of the terminal supporting multimodal scientific data input according to an embodiment of this application.
[0179] As shown in Figure 21, the terminal 2100 supporting multimodal scientific research data input includes an input device 2101. The input device 2101 may include at least one of the following: a keyboard, mouse, microphone, scanner, camera, handwriting tablet, drawing tablet, USB interface, network card, etc.
[0180] Referring again to Figure 21, the terminal 2100 supporting multimodal scientific data entry may further include: one or more memories 2102 and one or more processors 2103. The memories 2102 store a computer program that is run by the processors 2103. When the computer program is run by the processors 2103, the processors 2103 execute the multimodal scientific data recording method described above.
[0181] The terminal 2100 that supports multimodal scientific data entry can be part or all of a computer device that can implement multimodal scientific data recording methods through software, hardware, or a combination of software and hardware.
[0182] As shown in Figure 21, the terminal 2100 supporting multimodal scientific data entry includes one or more memories 2102, one or more processors 2103, a display (not shown), and a communication interface, etc. These components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). It should be noted that the components and structure of the terminal 2100 supporting multimodal scientific data entry shown in Figure 21 are exemplary and not limiting. The terminal 2100 supporting multimodal scientific data entry may also have other components and structures as needed.
[0183] Memory 2102 is used to store various data and executable program instructions generated during the operation of related methods, such as storing various application programs or algorithms that implement various specific functions. It may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0184] The processor 2103 may be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other processing units with data processing capabilities and / or instruction execution capabilities, and may be other components in the terminal 2100 that supports multimodal scientific data entry to perform the desired functions.
[0185] In one example, the terminal 2100 that supports multimodal scientific data entry also includes an output device that can output various information (such as images or sounds) to the outside (e.g., an operator), and may include one or more of a display device, a speaker, etc.
[0186] The communication interface can be any known communication protocol interface, such as a wired interface or a wireless interface. The communication interface may include one or more serial ports, USB interfaces, Ethernet ports, WiFi, wired networks, DVI interfaces, device integrated interconnect modules, or other suitable ports, interfaces, or connections.
[0187] Furthermore, according to embodiments of this application, a storage medium is also provided, on which program instructions are stored. When the program instructions are executed by a computer or processor, they are used to perform corresponding steps of the multimodal scientific data recording method of this application. The storage medium may, for example, include a memory card of a smartphone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disc read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.
[0188] Furthermore, embodiments of this application also provide a computer program product, which, when executed by a processor, implements the steps of the method described above.
[0189] The multimodal scientific data entry system, terminal, storage medium, and computer program product of this application embodiment have the same advantages as the aforementioned multimodal scientific data recording method because they can implement the aforementioned multimodal scientific data recording method.
[0190] Although exemplary embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above exemplary embodiments are merely illustrative and are not intended to limit the scope of this application. Various changes and modifications can be made therein by those skilled in the art without departing from the scope and spirit of this application. All such changes and modifications are intended to be included within the scope of this application as claimed in the appended claims.
[0191] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0192] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed.
[0193] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0194] Similarly, it should be understood that, in order to streamline this application and aid in understanding one or more of the various inventive aspects, features of this application may sometimes be grouped together in a single embodiment, figure, or description thereof in the description of exemplary embodiments of this application. However, this approach should not be construed as reflecting an intention that the claimed application requires more features than are expressly recited in each claim. Rather, as reflected in the corresponding claims, its inventive point lies in solving the corresponding technical problem with features fewer than all features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into that detailed description, wherein each claim itself is a separate embodiment of this application.
[0195] Those skilled in the art will understand that, apart from the mutual exclusion of features, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or apparatus so disclosed can be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.
[0196] Furthermore, those skilled in the art will understand that although some embodiments described herein include certain features but not others included in other embodiments, combinations of features from different embodiments are intended to be within the scope of this application and form different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.
[0197] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some modules according to the embodiments of this application. This application can also be implemented as an apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such an implementation of this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0198] It should be noted that the above embodiments are illustrative of this application and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In the unit claims enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0199] The above description is merely a specific embodiment or illustration of the embodiments of this application. The scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. The scope of protection of this application shall be determined by the scope of the claims.
Claims
1. A method for recording multimodal scientific research data, characterized in that, The method is applied to a multimodal scientific research data entry system, the multimodal scientific research data entry system including a backend module; the method includes: Based on the current research unit scheme, multimodal research data collected by the input device is received through the input interface; wherein, the multimodal research data includes at least one of the following data forms: text data, voice data, image data, and structured data; Based on the data field format specifications, the backend module is invoked to perform structured processing on the multimodal scientific research data, so as to convert the multimodal scientific research data into structured operation instructions; The structured operation instructions are converted into structured research records according to the research unit scheme.
2. The method according to claim 1, characterized in that, Based on the data field format specifications, the backend module is invoked to perform structured processing on the multimodal research data, so that the multimodal research data is converted into structured operation instructions, including: The multimodal scientific research data is identified and classified by the data modality recognition module to determine the type of the multimodal scientific research data; In the case where the multimodal research data is unstructured, based on the data type of the multimodal research data, the encoder in the backend module corresponding to the data type of the multimodal research data is invoked to convert the multimodal research data into the structured operation instructions.
3. The method according to claim 2, characterized in that, The multimodal research data is speech data; the method further includes: Invoke the speech encoder corresponding to the speech data so that the speech encoder converts the speech data into unstructured text data; Output the unstructured text data.
4. The method according to claim 2, characterized in that, The multimodal scientific research data is image data; the method further includes: The image encoder corresponding to the image data is invoked so that the image encoder converts the image data into unstructured text data based on the labeled industry knowledge dataset using text recognition algorithms and / or optical character recognition algorithms; Output the unstructured text data.
5. The method according to claim 1, characterized in that, in, The text data in question is unstructured text data.
6. The method according to any one of claims 3 to 5, characterized in that, The multimodal scientific research data entry system also includes an interactive front-end module; the interactive front-end module includes a first display area and a second display area; the method further includes: Send unstructured text data to the first display area so that the first display area displays the unstructured text data in the form of text blocks; and / or The system receives unstructured text commands input by the operator in the first display area and responds to the unstructured text commands.
7. The method according to claim 6, characterized in that, The multimodal research data is structured data; the method includes: The structured text data is sent to the second display area so that the second display area displays the structured text data; The data type and values of the structured data should follow the custom rules defined in the data field input box of the research unit scheme.
8. The method according to claim 6, characterized in that, The method further includes: Play the unstructured text data via voice interaction; and / or Receive voice commands from operators and respond to those voice commands.
9. The method according to claim 6, characterized in that, According to the research unit scheme, the unstructured text instructions are converted into structured operation instructions, including: Based on the prompts of the current research unit scheme, and using a preset algorithm to analyze and understand the unstructured text instructions, the unstructured text data is converted into corresponding structured operation instructions.
10. The method according to claim 9, characterized in that, Based on the prompts of the current research unit scheme, and using a preset algorithm to analyze and understand the unstructured text instructions, the unstructured text instructions are converted into corresponding structured operation instructions, further including: An embedded tool specific to the research unit scheme prompt word architecture is used to vectorize each knowledge point in each text block and store it in the research unit scheme data field input box in the form of key-value pairs for subsequent fast matching indexing.
11. The method according to claim 10, characterized in that, in, The research unit scheme prompts are pre-set in the backend module; and / or, the research unit scheme prompts are automatically created by the backend module based on the knowledge base corresponding to the current research unit scheme.
12. The method according to claim 7, characterized in that, According to the research unit scheme, the structured operation instructions are converted into structured research records, including: The structured data is stored in the research unit scheme data field input box and displayed in the second display area.
13. The method according to claim 12, characterized in that, The method further includes: based on data field format specifications, calling the backend module to perform structured processing on the multimodal research data, so as to convert the multimodal research data into structured operation instructions, including: The backend module is invoked, and its built-in algorithm is used to directly convert the multimodal scientific research data into structured data. The structured data is stored in the research unit scheme data field input box and displayed in the second display area.
14. The method according to claim 12 or 13, characterized in that, The method further includes: The backend module converts the structured operation instructions into code marked with the structured operation instructions. Execute the code marked with the structured operation instructions to form a structured scientific research experiment record.
15. The method according to claim 14, characterized in that, The research unit scheme data field input box is also used to store multiple data records; the method further includes: Based on the constraint relationships and / or data verification relationships between the multiple data records, the multiple data records are verified, and if there are logical errors in the multiple data records, a prompt indicating that there are logical errors is sent.
16. The method according to claim 14, characterized in that, The method further includes: Based on the storage instructions input by the operator, after storing the experimental records in the data according to the research unit plan, the text records and text summaries stored in the data field input box of the research unit plan are released; or, Based on the non-storage command input by the operator, the text record and text summary stored in the data field input box of the scientific research unit are directly released.
17. The method according to claim 16, characterized in that, The method further includes: Based on the management instructions input by the operator, the content stored in the input box of the research unit plan data field is managed. The management operations include at least one of the following: deleting data, storing data, and modifying data.
18. The method according to claim 6, characterized in that, The first display area and the second display area are each provided with scroll bars. When there is a lot of content displayed in the first display area and / or the second display area, the scroll bars will automatically scroll to display the latest content.
19. The method according to claim 1, characterized in that, The method further includes: Obtain instructions to record the experimental data in a predetermined national language; The structured operation instructions are translated into the predetermined national language and recorded in the experimental log.
20. The method according to claim 1, characterized in that, The input device includes at least one of the following: keyboard, mouse, microphone, scanner, and camera.
21. A multimodal scientific research data entry system, characterized in that, The system includes a backend module and an interactive frontend module; wherein... The backend module is used to receive multimodal research data collected by at least one input device through an input interface based on the current research unit scheme; perform structured processing on the multimodal research data according to the data field format specifications to convert the multimodal research data into structured operation instructions; and convert the structured operation instructions into structured research records according to the research unit scheme; wherein, the multimodal research data includes at least one of the following data forms: text data, voice data, image data, and structured data; The interactive front-end module includes a first display area and a second display area. The first display area is used to display and / or play unstructured data, and the second display area is used to display structured data.
22. A terminal supporting multimodal scientific research data input, characterized in that, The terminal includes: At least one input device, wherein the input device acquires multimodal scientific research data and sends the multimodal scientific research data to a processor for processing; the input device includes at least one of the following: a keyboard, a mouse, a microphone, a scanner, and a camera; The memory and the processor, wherein the memory stores a computer program that is executed by the processor, the computer program, when executed by the processor, causes the processor to perform the multimodal scientific data recording method as described in any one of claims 1 to 20.
23. A storage medium, characterized in that, The storage medium stores a computer program, which, when run by a processor, causes the processor to execute the multimodal scientific data recording method as described in any one of claims 1 to 20.
24. A computer program product, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method as described in any one of claims 1 to 20.