Multimodal scientific research data recording method, system, terminal and storage medium
Through the multimodal scientific research data entry system, the multimodal scientific research data is structured using large language models and multimodal large models, solving the problem of inefficiency of traditional experimental recording methods and achieving efficient and accurate scientific research data recording.
Patent Information
- Application Number
- CN202510086467.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-01-20
AI Technical Summary
Traditional experimental recording methods are inefficient, complex in operation, and intimate correlation and integration of multimodal data, which affects the comprehensiveness and accuracy of the recording.
Through the multimodal scientific research data entry system, the multimodal scientific research data is structured using large language models and multimodal large models, structured operation instructions are generated, data is automatically identified and recorded, and errors are corrected through dialogue interaction.
It improves the efficiency of experimental recording, ensures data accuracy and consistency, reduces manual input time, and adapts to diversified scientific research data needs.
Smart Images

Figure CN120011362B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data analysis technology, and in particular to a multimodal scientific research data recording method, system, terminal and storage medium. Background Art
[0002] While traditional experimental recording methods can import text data, voice data, handwritten records taken from photographs, and standardized data output from experimental instruments and equipment, these data still require researchers to manually input them or passively import them into the recording platform, resulting in inefficient and complex recording. This is especially true when the experiment is busy or requires rapid recording, as existing tools are not convenient enough to meet these needs.
[0003] Moreover, although it can support multiple modal data, the association and integration between different modal data are not tight enough, which affects the comprehensiveness and accuracy of experimental records. In addition, traditional experimental records are strictly written according to the template format and are not flexible enough. Summary of the Invention
[0004] This application is proposed in view of at least one of the above technical problems existing in the prior art. According to a first solution of this application, a multimodal scientific research data recording method is provided, which is applied to a multimodal scientific research data entry system, wherein the multimodal scientific research data entry system includes a backend module; the method includes:
[0005] Based on the current scientific research unit scheme, receiving multimodal scientific research data collected by an input device through an input interface; wherein the multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data, and structured data;
[0006] Based on the data field format specification, calling the back-end module to perform structured processing on the multimodal scientific research data, so as to convert the multimodal scientific research data into structured operation instructions;
[0007] The structured operation instructions are converted into structured scientific research records according to the scientific research unit plan.
[0008] The multimodal scientific research data recording method of the embodiment of the present application structures the multimodal scientific research data to form structured operation instructions, and converts the structured operation instructions into structured scientific research records according to the scientific research unit plan. This can reduce manual input time and improve experimental recording efficiency. It also automatically identifies and records multimodal scientific research data, and corrects errors through dialogue interaction, ensuring data accuracy and consistency.
[0009] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.
[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention, as claimed. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0012] Figure 1 A schematic flow chart showing a multimodal scientific research data recording method according to an embodiment of the present application;
[0013] Figure 2 A schematic flowchart of step S102 according to an embodiment of the present application is shown;
[0014] Figure 3 A schematic diagram illustrating a method for recording multimodal scientific research data using an encoder to pre-process multimodal scientific research data according to an embodiment of the present application is shown;
[0015] Figure 4 A schematic flowchart showing step S102 according to another embodiment of the present application is shown;
[0016] Figure 5 A schematic diagram illustrating a method for recording multimodal scientific research data without using an encoder to pre-process the multimodal scientific research data according to another embodiment of the present application;
[0017] Figure 6 A schematic flow chart illustrating displaying unstructured text data using a first display area according to an embodiment of the present application is shown;
[0018] Figure 7 A schematic diagram illustrating playing unstructured text data in a voice interactive manner according to an embodiment of the present application is shown;
[0019] Figure 8 A schematic flowchart of step S102 according to an embodiment of the present application is shown;
[0020] Figure 9 A schematic flowchart of step S801 according to an embodiment of the present application is shown;
[0021] Figure 10 A schematic diagram illustrating an operation instruction for converting text data into structured data fields based on a scientific research unit solution prompt word according to an embodiment of the present application;
[0022] Figure 11 A schematic diagram illustrating converting image data into structured data field operation instructions using an image encoder based on a scientific research unit solution prompt word according to an embodiment of the present application;
[0023] Figure 12 A schematic diagram showing a display interface according to an embodiment of the present application;
[0024] Figure 13 A schematic flowchart showing step S103 according to another embodiment of the present application is shown;
[0025] Figure 14 A schematic flow chart showing the process of converting a structured operation instruction into a code marked with a structured operation instruction according to an embodiment of the present application is shown;
[0026] Figure 15 A schematic flow chart illustrating verification of data records based on data verification relationships and / or data constraint relationships between / among multiple data fields according to an embodiment of the present application is shown;
[0027] Figure 16 A schematic flow chart illustrating permanent storage and release of data records according to an embodiment of the present application is shown;
[0028] Figure 17 A schematic flow chart illustrating an operation for managing the content of a scientific research unit program data field input box according to an embodiment of the present application is shown;
[0029] Figure 18 A schematic flow chart showing recording an experiment in a predetermined national language according to an embodiment of the present application is shown;
[0030] Figure 19 A schematic diagram illustrating the interaction between an operator, an AI interaction front end, a data recording end, an AI back end, and an AI database according to an embodiment of the present application is shown;
[0031] Figure 20 A schematic block diagram of a multi-modal scientific research data entry system according to an embodiment of the present application is shown;
[0032] Figure 21 A schematic block diagram of a terminal supporting multimodal scientific research data entry according to an embodiment of the present application is shown. DETAILED DESCRIPTION
[0033] To help those skilled in the art better understand the technical solutions in the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0034] Experimental records in the natural sciences often consist of data from various sources, modalities, and representations (e.g., handwritten notes, spreadsheets, instrument photos, etc.). Furthermore, in some specialized environments and scenarios, researchers need to be able to use free-form methods such as voice interaction and conversations in flexible scenarios to complete experimental records. However, sharing and reusing data from different modalities is difficult.
[0035] Based on at least one of the aforementioned technical problems, the present application provides a multimodal scientific research data recording method, which is applied to a multimodal scientific research data entry system, wherein the multimodal scientific research data entry system includes an interactive front-end module and a back-end module: the interactive front-end module is mainly used for human-computer interaction; the back-end module is mainly used for structured processing and storage of multimodal scientific research data. The method includes: based on the current scientific research unit scheme, receiving multimodal scientific research data collected by an input device through the input interface of the interactive front-end module; wherein the multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data and structured data; based on the scientific research unit scheme and data field format specifications, calling the back-end module to perform structured processing on the multimodal scientific research data, so that the multimodal scientific research data is converted into structured operation instructions; and converting the structured operation instructions into structured scientific research records in accordance with the scientific research unit scheme. The multimodal scientific research data recording method of the embodiment of the present application structures the multimodal scientific research data to form structured operation instructions, and converts the structured operation instructions into structured scientific research records according to the scientific research unit plan. This can reduce manual input time and improve experimental recording efficiency. It also automatically identifies and records multimodal scientific research data, and corrects errors through dialogue interaction, ensuring data accuracy and consistency.
[0036] The executor of the embodiment of the present application may be a multimodal scientific research data entry system. The multimodal scientific research data entry system may be an artificial intelligence (AI) system, and the AI system may be based on large language models (LLMs) or large multimodal models (LMMs). This article does not distinguish between the three. The AI system is capable of structured processing of multimodal scientific research data; it can not only automatically record experimental data, but also interact with operators, correct various numerical errors and logical errors, etc., to ensure data accuracy and consistency.
[0037] The large language model in this article refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. The large language model can handle a variety of natural language tasks, such as text classification, question and answer, dialogue, etc. At present, the large language model adopts a transformer (Transformer) architecture and pre-training objectives similar to the small model. The difference from the small model is the increase in model size, training data and computing resources. The large language model (LLM) of the embodiment of the present application may further include a variety of encoders corresponding to a variety of modalities, and each encoder can convert the data of the corresponding modality into text or its equivalent representation (such as token or vector embedding), so that the large language model has the ability to understand and process multimodal information. The large language model can thus process multimodal scientific research data and generate a series of executable structured operation instructions therefrom.
[0038] The multimodal big model in this article refers to the combination of big models with multimodal data such as text, images, audio, and video for training, thereby achieving the fusion and understanding of multimodal information in a unified representation space. With this fusion capability, the multimodal big model can not only process text-level semantics, but also conduct more in-depth analysis and reasoning on non-textual data such as images, audio, and video, thereby demonstrating more comprehensive intelligent processing capabilities in complex tasks such as cross-modal retrieval, multimodal question-answering, and image-text generation. Under the paradigm of the multimodal big model, the information contained in multiple modal inputs can be directly processed by the multimodal big model, which generates a series of executable structured operation instructions.
[0039] Figure 1 A schematic flow chart of a multimodal scientific research data recording method according to an embodiment of the present application is shown; Figure 1 As shown, the multimodal scientific research data recording method 100 according to an embodiment of the present application may include the following steps S101, S102, and S103:
[0040] In step S101 , based on the current research unit protocol, multimodal research data collected by an input device is received through an input interface of an interactive front-end module.
[0041] The multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data, and structured data. The multimodal scientific research data of the embodiment of the present application can provide operators with a more intuitive and detailed description of research activities.
[0042] Text data in multimodal scientific research data can be text data entered through an input interface by an operator in the form of a human-computer dialogue between an operator and a large language model. This text data can be text expressed in natural language. Voice data can be a voice message entered by an operator through an input interface, or it can be in the form of audio or recording. Image data can be an image captured of an experimental subject, a picture taken from a manually written text, an image of the equipment, or even a video. Structured data can be structured data generated by experimental instruments / equipment or derived from them. Structured data typically has a clearly defined data model and structure, can be stored in a relational database, and can be queried and manipulated using SQL (Structured Query Language). Structured data is characterized by its high degree of organization. Data is stored in rows and columns with a fixed format and length, and each field has a predefined data type, such as integer, string, or date. The advantages of structured data include ease of retrieval, update, and deletion, as well as data consistency and accuracy. Common structured data types include numbers, dates, times, phone numbers, and addresses. The structured data may also include structured data generated by an instrument application programming interface (API) / computational model, etc. In the embodiment of the present application, the operator may directly input the structured data through the interactive front-end module.
[0043] The input interface may be a human-machine interface defined by an operator, and multimodal scientific research data may be received through the input interface.
[0044] The input device includes at least one of the following: a keyboard, a mouse, a microphone, a scanner, a camera, a handwriting tablet, a drawing tablet, a USB interface, a network card, etc.
[0045] The research unit solution here essentially consists of components such as the research protocol, model, and assigner. In the research unit solution protocol, operators can customize various types of data fields and their corresponding data field format specifications (Data Field JSON Schema).
[0046] In step S102, based on the scientific research unit scheme and data field format specification, the back-end module is called to process the multimodal scientific research data so as to convert the multimodal scientific research data into structured operation instructions.
[0047] The structured operation instructions here refer to a series of executable single / multiple operation instructions.
[0048] In an embodiment of the present application, the multimodal scientific research data involved in the current scientific research unit plan can be sent to the back-end module for processing. The large language model in the embodiment of the present application can include a speech encoder, an image encoder, a structured data encoder, and other encoders. Each modality of data corresponds to its own encoder. For example, speech data corresponds to a speech encoder, and image data corresponds to a picture encoder.
[0049] It is worth noting that after parsing the various data, the various encoders mentioned above convert them into text data according to the research unit scheme. The large language model then processes the text data to obtain structured operation instructions. Furthermore, because the various encoders mentioned above parse the various data into text data, the text data in the multimodal research data does not need to be parsed by additional encoders. Instead, the text data can be sent directly to the large language model for processing.
[0050] In one embodiment of the present application, Figure 2 As shown, step S102 includes step S201, step S202 and step S203:
[0051] In step S201, the multimodal scientific research data is identified and classified by a data modality identification module to determine the type of the multimodal scientific research data;
[0052] In step S202, based on the data type of the multimodal scientific research data, an encoder and a large language model corresponding to the data type of the multimodal scientific research data in a backend module are called to convert the multimodal scientific research data into the structured operation instructions;
[0053] In step S203, the backend processing module converts the structured operation instructions into structured scientific research data.
[0054] The structured operation instructions of the embodiment of the present application can be composed of multiple operation instructions that can be directly executed, and when multiple operation instructions are executed, structured data can be generated. The structured data will be stored in the data field input box of each scientific research unit plan and eventually form an experimental record.
[0055] For example, the structured operation instructions expressed in a certain programming language are as follows:
[0056]
[0057] Since the structured research data in a research record can actually be stored as a series of key-value pairs under the research unit scheme paradigm, the structured research data generated based on the above operation instructions is:
[0058]
[0059] That is, the experimenter is updated to "Zhang San", the experimental temperature is updated to 25 degrees, and the experimental humidity is updated to 50%.
[0060] The core technical solutions of the embodiments of this application are to convert the data in the multimodal scientific research data input such as text / voice / picture into structured operation instructions, so that the system can insert the relevant information of the structured scientific research data contained therein into the corresponding data fields based on these structured operation instructions.
[0061] In one embodiment of the present application, the multimodal scientific research data can be structured according to the data type of the multimodal scientific research data. Figure 3 FIG2 is a schematic diagram showing how to process multimodal scientific research data using encoders of corresponding types based on the types of multimodal scientific research data according to an embodiment of the present application. The following examples illustrate the processing of different types of multimodal scientific research data.
[0062] In the first example, the multimodal scientific research data is speech data. First, a speech encoder corresponding to the speech data is called to enable the speech encoder to convert the speech data into text. Then, the text is output. For example, Figure 12 When the operator speaks the voice of "search PCR", the voice encoder converts the voice data into the text form of "search PCR" and presents it in the first display area.
[0063] In a second example, the multimodal scientific research data is image data. First, an image encoder corresponding to the image data is invoked. The image encoder converts the image data into text using a text recognition algorithm and / or an optical character recognition algorithm based on the annotated industry knowledge dataset. The text is then output.
[0064] Continue to combine Figure 3 For example, image data may include photos of handwritten text and photos of equipment. For handwritten text images, a text recognition algorithm (texttract algorithm) may be used to extract text from the handwritten text image and recognize the handwritten text as text. For experimental objects, a text recognition algorithm (texttract algorithm) and an optical character recognition algorithm (OCR technology) may be used to perform image analysis on the equipment image, extract the text and target object therein, and then generate text. For equipment images, an optical character recognition algorithm (OCR technology) may be used to perform image analysis on the equipment image, obtain the state of the equipment, and then generate text. For example, if it is found through image analysis that the lid of a certain device is not properly closed, the text "The lid of a certain device is not properly closed" may be recorded in the experimental record in the form of text. This method not only directly converts image data into text in the terminal, but also provides prompts for irregular experimental operations to improve the accuracy of the experiment.
[0065] Combine Figure 12 For example, before a formal experiment, an operator needs to record the experimental environment. For example, a photo of a piece of experimental equipment is taken and then input into the AI system. The AI system's backend analyzes the photo and discovers that the lid of a piece of experimental equipment is not properly closed. The AI system records this information and displays it in the second display area 1202 as a text annotation, i.e., "The lid of a piece of equipment is not properly closed."
[0066] It is worth noting that video data is composed of multiple frames of pictures. Video data belongs to image data, and the processing of video data is consistent with the image data processing process.
[0067] Among them, image data is an important unstructured data source. In scientific experiments, it is often necessary to take pictures of the experimental process, and operators sometimes need to handwrite experimental records, etc. The embodiment of the present application will call a large language model to identify image-based scientific research data and enter it into the scientific research unit program data field input box (RU Data Field InputBox) for storage. The processing of image data is mainly divided into two cases: first, in the initialization stage, image information is used for initialization to identify and record the scientific research data contained in the image information; second, consistent with the above-mentioned scenario in the experimental record of text data, the input of image data is supported in the dialogue.
[0068] In the fourth example, the multimodal scientific research data is structured data. The present application embodiment also supports experimental records of structured markup documents such as Word, JSON, Markdown, etc., and these structured documents can be identified and entered using large language models and structured parsing tools.
[0069] In an embodiment of the present application, the operator can manually fill in the structured data in the scientific research unit program data field input box according to the current scientific research unit program. For example, the operator manually fills in the experimental data such as the experimenter, experiment name, experiment date, experiment temperature, experimental object, and change value of the experimental object in the scientific research unit program data field input box. It is worth noting that when manually filling in the structured data, it can be directly filled in the scientific research unit program data field input box in the second display area so that the second display area displays the structured data. Among them, the data type and value of the structured data should follow the rules customized by the scientific research unit program data field input box.
[0070] Regarding the experimental records of the outputs of instruments and computing models, since the instruments and computing models widely used in modern scientific experiments can automatically generate a large amount of data, such as experimental measurement values, instrument status, statistical results, etc., the embodiments of the present application can enable the large language model to connect to the APIs of such instruments, equipment and models to obtain input and output information and perform experimental records.
[0071] In addition, the structured operation instructions generated by the large language model are text. When acquiring data such as voice data, image data, and structured data, these data need to be converted into text (e.g., voice, image, video, etc. data encoded in Base64 string format; or the voice, image, video, etc. data are uploaded to a local database / cloud database / local area network / internet to obtain the corresponding text ID / URL / URI). The text is then stored in the research unit plan data field input box and displayed on the display interface. If the structured operation instructions generated by the large language model contain voice data, these voice data can be stored in the recording slot and played when a play command is received. If the structured operation instructions generated by the large language model contain image data, these image data can be stored in the recording slot and displayed on the display interface or preview window when a display command is received. If the structured operation instructions generated by the large language model contain video data, these video data can be stored in the recording slot and played through a video player when a play command is received. If the structured operation instructions generated by the large language model contain data in the form of structured documents, these data in the form of structured documents can be stored in the record slot and automatically parsed and presented according to the preset format requirements when receiving the parsing or display instructions.
[0072] The embodiments of the present application can realize intelligent experimental recording on a universal and extensible platform, that is, automatically identifying and recording standardized data such as text data, voice data (including voice commands), handwritten text and pictures, and output of experimental instruments and equipment.
[0073] In another embodiment of the present application, Figure 4 As shown, step S102, based on the data field format specification, calls the back-end module (including the multimodal large model) to perform structured processing on the multimodal scientific research data, so as to convert the multimodal scientific research data into structured operation instructions (single / multiple operation instructions), including steps S401 and S402:
[0074] In step S401, the backend module is called and the built-in algorithm of the backend module is used to directly convert the multimodal scientific research data into structured operation instructions;
[0075] In step S402, the structured scientific research data contained in the structured operation instruction is stored in the scientific research unit plan data field input box and displayed in the second display area.
[0076] like Figure 5 As shown, after receiving the multimodal scientific research data (text data, voice data, image data, instrument API / computational model and other data), the input device sends the multimodal scientific research data directly to the multimodal large model, so that the multimodal large model uses the built-in algorithm for processing. For example, the multimodal large model can use a deep learning algorithm to analyze and understand the multimodal scientific research data, generate a plurality of structured operation instructions, and automatically execute the plurality of structured operation instructions to generate corresponding structured text data. These structured text data will be stored in the scientific research unit program data field input box. In addition, after analyzing and understanding the multimodal scientific research data, the multimodal large model can output unstructured text data and display the structured text data on the display interface (such as the first display area) so that the user can intuitively understand the content of the experimental record.
[0077] In one example, combining Figure 12 , Figure 12 The second display area 1202 in the diagram displays the structured text data generated by executing the structured operation instructions. This structured text data will be presented in the final experiment record. Users can directly enter the corresponding content, i.e., structured text data, in the research unit plan data field input box corresponding to data fields such as "Experimenter," "Experiment Date," "Experiment Purpose," "Temperature," and "Humidity."
[0078] In another example, continue with Figure 12 , Figure 12The first display area 1201 in the figure displays unstructured text data, which is obtained by the AI system through analysis and understanding of multimodal data (such as voice, images, text, etc.). This unstructured text data needs to be processed again before it can be presented as structured text data such as that displayed in the second display area 1202. For example, the operator asks the question "Please change the temperature to 45 degrees" through voice. After analyzing and understanding and executing the operation, the AI system responds with a voice message "The temperature has been changed to 45 degrees."
[0079] That is, in this example, the AI system converts the multimodal data into unstructured text data (e.g., the answer generated by the AI based on the user's question) and presents it in the first display area 1201. It is worth noting that the user can enter a question in the form of text data. For example, if the user enters the question "Search for PCR" in the question input box of the first display area 1201, the AI system will not be able to find PCR after searching, and will display "Unable to search for PCR" as text data.
[0080] It is worth noting that the current scientific research unit solution not only performs intuitive language conversion for the input multimodal scientific research data, but also can understand the input data based on algorithms such as deep learning algorithms. Problems that arise in the actual process can be recorded and displayed in the final experimental record. For example, when a photo (or video) of a certain device is parsed and it is found that the cover of a certain device is not properly closed, it will be recorded in the data field input box of the scientific research unit solution, "The cover of a certain device is not properly closed". The embodiment of the present application can parse the photo, obtain the state of the object in the photo, and record the state of the object in the data field input box of the scientific research unit solution to facilitate the operator to make timely adjustments.
[0081] The AI system of the embodiment of the present application can better integrate, analyze and understand multimodal data. For example, for the image and voice data obtained during the experiment through the camera and microphone, the image encoder and voice encoder can be used to convert them into text and then use the large language model for analysis and understanding, or directly use the multimodal large model for analysis and understanding to generate corresponding experimental records. In addition, since the understanding of the large language model / multimodal large model has high accuracy and flexibility, it can answer questions related to the experiment and assist operators in recording experiments.
[0082] like Figure 6 As shown, the method further includes step S601 and step S602:
[0083] In step S601, unstructured text data is sent to the first display area, so that the first display area displays the unstructured text data in the form of text blocks; and / or
[0084] In step S602, an unstructured text instruction input by an operator in the first display area is received, and a response is made to the unstructured text instruction.
[0085] In one embodiment of the present application, the multimodal scientific research data entry system further includes an interactive front-end module. Figure 12 As shown, the display interface 1200 of the interactive front-end module includes a first display area and a second display area. The first display area is used to display human-computer question-and-answer content in text form, or to display voice data (parsed into text data by a voice encoder), image data (parsed into text data by an image encoder), or structured data (parsed into text data by a structure encoder). Please refer to the following for a detailed introduction to the first and second display areas.
[0086] For example, natural language processing (NLP) technology is used to perform speech-to-text operations on the voice interaction data to generate text corresponding to the voice; the voice interaction data may include voice instructions, requiring the AI system to record scientific research data, and / or answer questions related to the scientific research unit plan and record key information during the experiment.
[0087] In some embodiments, the content displayed in the first display area can be played out in the form of voice at the same time as it is displayed, so that the operator can know the displayed content and the feedback of the AI system without having to go to the front of the display interface. In addition, the user can also directly use voice to give command feedback to the broadcast voice, so that the AI system can further process and provide feedback based on the user's further voice commands. Figure 7 As shown, the method further includes step S701 and step S702:
[0088] In step S701, the unstructured text data is played in a voice interactive manner; and / or
[0089] In step S702, a voice instruction from an operator is received and a response is made to the voice instruction.
[0090] During multiple rounds of human-computer question-answering (voice interaction), the large language model can use natural language processing technology to perform voice-to-text conversion and text summary generation on the voice interaction data. The operator can input experimental records through voice, and the large language model will convert the voice into text and store it. Through the NLP-based dialogue system, it can interact with the operator, answer questions related to the experiment, and assist in recording key information during the experiment.
[0091] In some embodiments, the operator can also issue some operation instructions to the large language model. After executing these operation instructions, the AI system can provide feedback to the operator on the execution status. For example, when the operator starts the experiment, he says "please initialize". The conversation module (for example, a dialogue robot) in the AI system obtains the initialization status of each experimental instrument or equipment based on the generated text summary "initialize". When all experimental instruments or equipment have completed the initialization operation, the voice answer is "initialization completed". For another example, the operator says "such and such experiment, Zhang San", the AI system can understand "Zhang San" as the experimental operator or experimental recorder according to the natural language processing algorithm, and then record "Operator: Zhang San" or "Experimental recorder: Zhang San" in the experimental record.
[0092] In addition, based on operator-defined rules, the AI system can be configured to maintain a human-computer question-and-answer process throughout the entire experimental process, so as to record data generated by multiple experimental steps in a single experiment or multiple scientific research data from multiple experiments.
[0093] In addition, the AI system can not only answer the operator's questions, but also search based on the keywords in the questions. In the embodiment of the present application, the operator can issue instructions to the AI system in the form of multiple rounds of dialogue based on the experimental process data displayed on the display interface to perform operations such as modifying data, correcting data, deleting data, storing data, and updating data.
[0094] For example, an operator can say "Help me search for PCR" and search for "PCR" as a field name. If "PCR" is not found, the operator can provide text / voice feedback such as "PCR not found, please try again". The operator can set the search count to 3 times before responding or stop searching if no results are found after 3 searches. The embodiment of the present application can perform semantic analysis on voice data, execute instructions based on their meaning, and then convert the execution results into text and provide feedback to the operator in the form of voice data.
[0095] In addition, the operator can conduct multiple rounds of question and answer sessions with the AI system at any time throughout the experiment, and the question and answer process and results can be stored and displayed in the scientific research unit program data field input box. In order to address the problem that operators cannot free their hands to record scientific research data in certain specific scenarios, the embodiment of the present application can also record the experimental process based on voice interaction, using the AI system to convert the operator's voice input and model output into the scenes in the experimental record of text data mentioned above, to complete the experimental record based on voice dialogue instructions.
[0096] For example, the user can give the system a command through voice: "Please start recording. Record the experimenter as Zhang San". The system responds: "The experimenter has been recorded as Zhang San" and announces it in voice form. After listening to the voice, the user can continue to give a command through voice: "Please record the experiment date as January 1, 2024". The system immediately responds: "The experiment date has been recorded as January 1, 2024". In this way, the user can directly record scientific research through the voice control system without approaching the recording interactive terminal or even checking the interactive interface. In this way, it can effectively cope with special experimental scenarios where both hands cannot be freed for keyboard operations. For example, when the user is conducting a chemical experiment in a glove box, he needs to continue to operate inside the glove box and cannot use the keyboard to enter data.
[0097] Therefore, the embodiments of the present application can adapt to the diverse needs of scientific research data and effectively improve the convenience and accuracy of scientific researchers' operations during the recording process.
[0098] In step S102, the unstructured text instructions may be converted into structured operation instructions according to the scientific research unit plan.
[0099] like Figure 8 As shown, step S102 converts the unstructured text instructions into structured operation instructions according to the scientific research unit plan, including step S801:
[0100] In step S801, based on the current scientific research unit program prompt words, a preset algorithm is used to analyze and understand the unstructured text instructions to convert the unstructured text data into corresponding structured operation instructions.
[0101] For example, a user issues the following unstructured text instruction:
[0102] "Please set the temperature to 25°C."
[0103] Through S801, it can be converted into the following structured operation instructions:
[0104]
[0105] In one embodiment of the present application, the encoder converts the multimodal scientific research data into the text based on the current scientific research unit scheme prompt word.
[0106] The research unit plan prompt can be automatically generated based on the current research unit plan, or it can be preset by the operator or other operators based on the current research unit plan. For example, the prompt "Temperature range is 20℃-50℃" can be displayed below the experimental temperature field. If the obtained temperature is not within this temperature range, a temperature error prompt will be displayed.
[0107] like Figure 9 As shown, step S801 is based on the current scientific research unit program prompt word, and uses a preset algorithm to analyze and understand the unstructured text data to convert the unstructured text data into corresponding structured operation instructions, further including step S901:
[0108] In step S901, an embedded tool specific to the scientific research unit scheme prompt word architecture is used to vectorize each knowledge point in each text block, and store it in the form of key-value pairs in the scientific research unit scheme data field input box for subsequent fast matching indexing.
[0109] The scientific research unit solution prompt words are preset in the back-end module; and / or the scientific research unit solution prompt words are automatically created by the back-end module according to the knowledge base corresponding to the current scientific research unit solution.
[0110] In one embodiment of the present application, Figure 10 and Figure 11 This article introduces how to convert the unstructured text data into corresponding structured operation instructions based on the scientific research unit program prompts. Figure 10 A schematic diagram illustrating converting text data into structured operation instructions based on scientific research unit solution prompts according to an embodiment of the present application is shown; Figure 11 A schematic diagram of converting image data into structured operation instructions according to the scientific research unit scheme prompt words in an embodiment of the present application is shown. The scientific research unit scheme prompt words are created by the AI system based on the current scientific research unit scheme. The AI system can also construct the scientific research unit scheme background prompt words by itself based on the background knowledge of the experiment. Background knowledge may include basic information defined by the operator, working environment, experimental tasks, experimental purposes, etc. For example, before conducting a medical experiment to treat lung disease, basic information such as room temperature and whether working in a sterile environment is input into the AI system, from which the current scientific research unit scheme background prompt words can be extracted and added to the original scientific research unit scheme prompt words.
[0111] like Figure 12As shown, it is a schematic diagram of the display interface 1200 of the multimodal scientific research data recording method of an embodiment of the present application. The display interface 1200 implemented in the present application includes a first display area 1201 and a second display area 1202. The first display area 1201 is located on the left side, and is used to display the content of the human-computer question and answer; the second display area 1202 is located on the right side, and is used to display the content temporarily stored in the scientific research unit program data field input box. In addition, the first display area 1201 and the second display area 1202 are respectively provided with scroll bars to facilitate the user to manually / automatically scroll the scroll bar when the human-computer question and answer content and the scientific research unit program data field input box and their contents are large, so as to display more content.
[0112] like Figure 13 As shown, step S103 converts the structured operation instruction into a structured scientific research record according to the scientific research unit plan, and further includes step S1301:
[0113] In step S1301, the structured data is stored in the scientific research unit plan data field input box and displayed in the second display area.
[0114] Combine Figure 12 , since different data types (such as strings, integers, floating-point numbers, Boolean values, dates, enumeration values, etc.) are defined for different fields in the data field format specification, in this case, different scientific research unit scheme data field input boxes can be generated for the data fields in the second display area, and corresponding interactive controls can be generated for the scientific research unit scheme data field input boxes based on the data types specially annotated for each field in the data field format specification. For example, annotations based on data types can be provided for the scientific research unit scheme data field input boxes based on the data type rules of the data field format specification. The embodiment of the present application can help operators focus their time and energy on the definition and development of substantive content such as scientific research protocols and data fields in scientific research schemes through this display method, without having to worry about issues such as experimental record interfaces, scientific research data storage structures and methods, thereby enabling scientists to efficiently design high-quality scientific research schemes that meet actual scientific research needs in an operator-friendly manner during daily scientific research activities, and use them for scientific research data recording.
[0115] Various scientific research data are stored in the database in the form of fields. It is worth noting that in the embodiment of the present application, since the AI system can directly obtain the scientific research data output by various experimental instruments or equipment, all the data in the various fields on the right side of the display interface can be directly displayed without manual filling, which improves efficiency compared to the traditional manual filling method. Moreover, the embodiment of the present application displays all the data in a visual form, which is convenient for operators to see the scientific research data at any time. In the case of a large amount of data on the right side, the scientific research data can be automatically sorted and the automatic scrolling interface display can be set, which also greatly improves the efficiency of experimental recording.
[0116] As shown in the figure, Figure 14 As shown, the method further includes step S1401 and step S1402:
[0117] In step S1401, the structured operation instruction is converted into a code marked with the structured operation instruction by the back-end module;
[0118] In step S1402, the code marked with the structured operation instruction is executed to form a structured scientific research record.
[0119] The code marked by the structured operation instruction can be represented as a plurality of operation instructions expressed in a certain machine language and capable of being automatically executed by the terminal.
[0120] In one embodiment of the present application, the scientific research unit plan data field input box is further used to store data corresponding to multiple data fields. The data corresponding to multiple data fields refers to data in corresponding data fields extracted from multimodal scientific research data such as text / voice / image input by the user.
[0121] For example, if the user says the temperature is set to 45 degrees and the humidity is 50%, then the corresponding processed structured scientific research data is
[0122] {
[0123] "temperature":25.0,
[0124] "humidity":50.0
[0125] }
[0126] The 25.0 and 50.0 here are the corresponding data in the data field.
[0127] like Figure 15 As shown, the method further includes step S1501:
[0128] In step S1501, based on the data verification relationship and / or constraint relationship between / among the multiple data fields, the data corresponding to the multiple data fields are verified, and if there is a logical error in the data corresponding to the multiple fields, a prompt of the existence of a logical error is sent.
[0129] Specifically, the type constraint includes constraining the corresponding data field to use predefined multimodal scientific research data during data entry, wherein the predefined multimodal scientific research data includes one or more of text data, image data, video data, audio data, and text data. In some embodiments, the operator can further define the data type (e.g., various numerical types, time types, etc.) of the data field previously defined in the scientific research protocol. For example, if the solvent_volume data field is defined as a floating point number, then when the operator enters non-floating point data (e.g., a string of letters) in this field, the system will indicate a type error.
[0130] The numerical verification relationship includes constraining the corresponding data field to follow a specified pattern and / or not exceed a preset value range when entering data. The operator can further define validation rules for the data fields defined in the scientific research agreement. For example, further validation rules are added for the solvent_volume data field and its type constraint (floating point number) to ensure that the floating point number filled in the data field must be greater than zero. In this case, if the operator enters a negative floating point number, the system will display a numerical verification error. In another example, the so-called following a specified pattern can be, for example, using RegExp regular expressions RE (Regular Expression) to constrain the composition rules and patterns of the character strings in the field. For example, the value recorded in a field related to an email address must contain and only have one "@" symbol, and so on. In other examples, other patterns and value range constraints can also be set, which are not listed here one by one.
[0131] The combined validation relationship involves constraining the types and / or values of each data field to meet a predetermined constraint. For example, in an experiment, a constraint may be that when the required temperature is greater than 40 degrees, the required humidity must be less than 20%. Thus, the temperature-related values and the humidity-related values form a combined validation relationship and are mutually dependent.
[0132] This allows for rapid identification of abnormal scientific research data. If any validation relationship fails, the reason for the failure will be displayed, and the operator can correct the invalid data field values one by one according to the error message until the record is valid. Through this validation process, the platform ensures that the data entered by the operator meets the requirements.
[0133] In some other embodiments, if the operator defines multiple data fields with dependency and assignment relationships in the scientific research protocol, then these relationships can be further customized using an assigner. Specifically, when the operator designs a scientific research unit plan for the corresponding discipline based on the scientific research node design environment, it can also include an assigner for the data field, and the assigner is used to assign values to the data field based on the data field dependency graph and the assignment rules. In some embodiments, the data field dependency graph is a single-level or multi-level directed acyclic graph, and the assignment relationship between the defined data fields is single dependency or multiple dependency, and each data field is assigned by at most one assigner. An assigner can take one or more upstream data fields as dependencies, and at the same time, a data field can actually serve as a dependency of one or more data fields, but for a specific data field, the way of determining its field value should be unique. Just as an example, if two data fields solvent_volume and solvent_volume_2 are defined in the research protocol, and the operator uses the assigner to define that solvent_volume_2 is always twice the value of solvent_volume, in this case, every time a new value is entered for solvent_volume, the system will automatically set solvent_volume_2 to twice the solvent volume value. For example, if the value of solvent_volume is 5, the system will automatically assign a value of 10 to solvent_volume_2. This can greatly improve the efficiency and accuracy of scientific data recording.
[0134] By defining the model of the scientific research unit plan, the operator's input can be dynamically verified to ensure the accuracy of data entry; through the definition of the assigner, multi-level, multi-dependent field dependencies can be automatically calculated according to the value of a certain input field, thereby ensuring that even if there are complex dependencies between multiple data fields, they can be efficiently entered and run correctly. This can significantly promote the laboratory's scientific research data, including the electronic management of data and orders for outsourced experiments, and realize efficient retrieval of scientific research plans, scientific research data and other related content.
[0135] For example, combined with Figure 12, the temperature data and humidity data obtained by the large language model (LLM) are 45 degrees Celsius and 70% respectively, while when the temperature is 45 degrees Celsius, the humidity cannot be 70%. At this time, a prompt can be issued in the display interface, for example, the humidity data can be highlighted to attract the attention of the operator; a sound prompt can also be issued using a buzzer, etc.; or it can be displayed in the form of text below the humidity field to prompt the operator. The embodiment of the present application can be set to verify the data immediately when each data is obtained, and after all the data are obtained, it is verified again based on the combination verification relationship between the related data to ensure the accuracy of the experiment and the experimental record.
[0136] In the embodiment of the present application, the experimental records of different modalities are converted into data forms (e.g., text) that can be recognized by a large language model (LLM) through an encoder of the corresponding modality. Subsequently, the powerful dialogue and task understanding capabilities of the current large language model (LLM) are used to update and record the experimental records of different modalities. In addition, in the process of dialogue with the large model, scientific research data with errors in recognition of other modalities can also be corrected through interactive input such as text and voice. While completing the recognition and recording of scientific research data, high-quality data for training and fine-tuning the multimodal scientific research large model is collected.
[0137] In one embodiment of the present application, Figure 16 As shown, the method further includes step S1601 and step S1602:
[0138] In step S1601, based on the storage instruction input by the operator, after the structured scientific research record is stored in the database according to the scientific research unit plan, the data corresponding to the data field stored in the data field input box of the scientific research unit plan is released; or,
[0139] In step S1602, based on the non-storage instruction input by the operator, the data corresponding to the data field stored in the scientific research unit plan data field input box is directly released.
[0140] In an embodiment of the present application, the data in the scientific research unit program data field input box is not permanently stored data, but temporarily stores various data in the scientific research / experimental process. After the final structured scientific research record is generated, the data in the scientific research unit program data field input box will be cleared or migrated to other storage devices. For example, the multimodal scientific research data entry system is installed on a local terminal, and the data stored in the scientific research unit program data field input box is correspondingly stored in the cache of the local terminal. When it is necessary to store it as the final structured scientific research record, the operator clicks the "Submit" button on the display interface, and the current structured scientific research record will be stored in the local database. For another example, the multimodal scientific research data entry system is installed on a cloud server, and the data stored in the scientific research unit program data field input box is correspondingly stored in the cloud server. When it is necessary to store it as the final structured scientific research record, the operator clicks the "Submit" button on the display interface, and the current structured scientific research record will be stored in the cloud database.
[0141] It is worth noting that when the structured scientific research records are stored in the database, a structured storage scheme will be automatically generated for the scientific research unit scheme based on the data field format specification, so that the operator can submit the structured scientific research records based on the scientific research unit scheme for storage, and each data field should comply with the constraints and data verification relationship of the model. In an embodiment of the present application, regardless of whether the data field defined by the operator is a simple text or a more complex multimodal data, the large language model can automatically generate a corresponding data structure for it and ensure the consistency and integrity of the data during the entry and storage process. Since the automatically generated data structure is based on standardized protocols and definitions, scientific research data can be easily shared globally. This automation of structured storage not only simplifies the work of scientific researchers, but also improves the reproducibility of scientific research data and the ability to collaborate across laboratories. In addition, through the automatically generated data structure, scientific researchers can focus on experiments and data records without worrying about underlying data management issues. This feature greatly reduces the data management burden of scientific researchers and improves the efficiency of scientific research work.
[0142] In one embodiment of the present application, Figure 17 As shown, the method further includes step S1701:
[0143] In step S1701, based on the management instruction input by the operator, the content stored in the input box of the scientific research unit plan data field is managed.
[0144] The management operation includes at least one of the following operations: deleting data, storing data, and modifying data.
[0145] For example, the operator manually deletes, modifies, and performs other operations on the experimental data in the scientific research unit program data field input box.
[0146] For example, when an operator discovers a problem with the temperature data in a set of scientific research data, he or she may say, "Correct the third set of temperatures to 45°C." The AI system receives the voice data and can directly process it using a built-in algorithm, converting it into structured operation instructions:
[0147]
[0148] Then, based on the above structured operation instructions, the AI system will fill in the structured scientific research data (the corresponding scientific research unit plan data field is "temperature" and the value is "45℃") into the corresponding scientific research unit plan data field input box; or, the large language model converts the voice data into unstructured text data through the voice encoder, and then converts the unstructured text data into structured text data, that is, the scientific research unit plan data field is "temperature" and the value is "45℃". Then, the structured data is filled in the corresponding scientific research unit plan data field input box.
[0149] For another example, when the temperature of the experiment is entered incorrectly and needs to be modified, the operator says "Change the third set of temperatures to 45°C". After the AI system changes the value in the scientific research unit plan data field input box of the temperature field to 45°C, the AI system can also provide feedback, such as feeding back the text expressed in natural language "The third set of temperatures has been corrected to 45°C", or further broadcasting the text through voice, so that the operator can obtain feedback from the AI system through voice without directly viewing the first display area.
[0150] In one embodiment of the present application, Figure 18 As shown, the method further includes step S1801 and step S1802:
[0151] In step S1801, an instruction to record the experimental record in a predetermined national language is obtained;
[0152] In step S1802, the structured operation instructions are translated into the predetermined national language and recorded in the experimental record.
[0153] For example, if the operator's native language is Chinese and the desired predefined national language is English, the user can interact with the AI system in the language they are proficient in. For example, they can issue an operation instruction in Chinese: "Please record the weather as sunny" (the research unit plan contains a data field with the ID "weather"). The AI system can then automatically understand the user's operation instruction and generate the following structured operation instruction:
[0154] {
[0155] "operation":"update",
[0156] "field_id":"weather",
[0157] "field_value":"Sunny"
[0158] }
[0159] Since the operator's predetermined national language is English, the AI system can further translate the information in the above operation instructions into English ("sunny" is translated into "Sunny"), and fill in "Sunny" in the data field input box corresponding to "weather".
[0160] The embodiment of the present application can support the input of multimodal scientific research data in the question-answering process, and conduct the question-answering process based on non-text modal data such as pictures, videos or voices, thereby expanding the use scenarios of the question-answering process. In addition, by expressing different forms of information through different modal data, the expression of information is richer, thereby being able to provide richer information to the large model, helping to improve the accuracy of the large model's understanding of the input content, thereby generating more accurate reply content and improving the accuracy of the reply content. In the question-answering process, when the conversation content includes non-text content, the method of the embodiment of the present application performs multimodal intent recognition on the conversation content and the non-text content. The multimodal intent recognition result indicates the content relevance between the conversation content and the non-text content, so that the corresponding target reply model can be selected based on the content relevance, thereby improving the accuracy of the selected reply model, and then improving the accuracy of the reply content.
[0161] The following combination Figure 19 The present invention will be described in detail again by taking the AI model as an example.
[0162] like Figure 19 The figure shows a schematic diagram of the interaction between the operator, the AI interaction front end, the data recording end, the AI back end and the AI database.
[0163] Depend on Figure 19 As can be seen, the data in this embodiment of the application primarily interacts between the operator, the AI interaction front-end, the data recording end, the AI back-end, and the AI database. The data entry method can be divided into three main stages: the first stage is the data processing stage, the second stage is the output stage, and the third stage is the stage of storing the operator / AI interaction information.
[0164] In the first stage, the input multimodal scientific research data is sent to the AI backend for processing to generate structured operation instructions, that is, multiple operation instructions, and execute multiple operation instructions to obtain data records.
[0165] Here are the steps:
[0166] Step 1: Input the relevant information of the data fields (DFs) to be recorded in the AI interaction front end in the form of dialogue / voice / picture, etc.
[0167] Step 2: The AI interaction front-end sends the operator's instructions to the AI back-end.
[0168] Step 3: Based on the scientific research unit plan and data field format specifications (data field JSON Schema), AI processes the operator's front-end input information into structured operation instructions in the form of a list on the data record side.
[0169] Step 4: The AI backend sends the operation instructions to the data recording end.
[0170] In step 5a, the data recording end processes each operation instruction one by one and generates Acknowledge information corresponding to each operation instruction.
[0171] Step 5b: Every time the data recording terminal successfully processes an operation instruction, the data field input box filled in by the operation instruction and the filled value are reflected in real time on the data recording interface.
[0172] In the second stage, AI can also generate AI second output. The steps are as follows:
[0173] In step 6a, the data recording end transmits the tabular Acknowledge information to the AI backend.
[0174] In step 6b, the AI backend generates the AI second output, i.e., the reply to the operator, based on a. the operator's original input, b. the AI first output (operation instruction list), and c. the data recording end response (Acknowledge information list).
[0175] In step 6c, the AI backend sends the AI second output to the AI interaction frontend.
[0176] Step 6d: The AI interactive front-end displays the information from the previous step to the operator (usually in the form of a dialogue answer)
[0177] In the third stage, the data records can be saved to the database, and the content of the multi-round human-computer question and answer can also be saved to the database, where the database can be a local database or a cloud database. The steps are as follows:
[0178] Step 7: The AI backend stores information such as a. operator original input, b. AI first output (operation instruction list), c. data recording end response (Acknowledge information list), and d. AI second output into the AI database.
[0179] In an embodiment of the present application, first obtain a handwritten text photo and a device photo (in other embodiments, it can be a video), which can be manually input into the execution subject (for example, a computer) by the operator. Based on the scientific research unit plan prompt words corresponding to the specific scientific research unit plan, the image encoder converts the handwritten text photo and the device photo into corresponding descriptive text. A dialogue robot based on the AI system (the dialogue robot can be part of the AI system) outputs scientific research data. These scientific research data are stored in a temporary scientific research unit plan data field input box to facilitate temporary recording and modification of experimental results. After the operator confirms the results, the experimental record results of this experiment will be stored in the data end of the system.
[0180] Using the scientific research unit program data field input box can flexibly control the output format of the experimental record, and enable the model to modify and manually correct the recorded errors. These manual correction operations will be collected as manual feedback data to provide an instruction data basis for the subsequent training and fine-tuning of the unified joint representation multimodal model.
[0181] It is worth noting that the entire process, from inputting handwritten text photos and device photos into the image encoder (part of the large language model) to the large language model-based conversational robot (also part of the large language model (LLM)) storing the data in the recording slot, can be completed by the large language model (all the functions mentioned in this application can also be achieved by using the multimodal large model). This can help operators record scientific research data more quickly and reduce the workload of manual input.
[0182] The multimodal scientific research data recording method of the embodiment of the present application can be applied to technical fields such as medical records, legal document management, and educational data analysis that require AI system recording and intelligent management. As an example, it can be directly applied to electronic medical record scenarios. Similar to the scientific research described in other embodiments of the present application, in the medical / hospital / clinical scenario, each institution / department has the need to record medical data (such as electronic medical records), and the content and type specifications of the data required to be recorded in different departments, different diseases, and different medical scenarios are not the same. It can be imagined that if each records in a customized way, it will not only be inefficient, but also not conducive to sharing and experience accumulation. With the help of the scientific research activity management and application platform in the embodiment of the present application, operators in different professional fields can customize the protocols, models, and evaluators of medical records according to the scientific research unit grammar, tailor the medical record method of their profession according to specific needs, and the customized scientific research unit scheme can be widely shared in different departments, so that a one-time design can be achieved and applied throughout the hospital / across hospitals to achieve unification and standardization of specific medical records. In addition, when recording clinical medical data, it often occurs in doctor-patient consultation scenarios, where doctors may not be able to free their hands to operate computers or other equipment or handwrite clinical medical data during the consultation process. The multimodal data recording solution (for example, based on voice) proposed in this application provides a convenient solution to the recording needs of this scenario. In summary, this application has huge application potential and significant economic benefits in multiple industries.
[0183] The multimodal scientific research data recording method of the embodiment of the present application converts the multimodal scientific research data of the current scientific research unit plan into structured operation instructions, and then combines at least one round of question-and-answer process with the operator to convert the structured operation instructions into structured scientific research records according to the scientific research unit plan. This can reduce manual input time and improve experimental recording efficiency. In addition, it uses an AI system to automatically identify and record multimodal scientific research data, and correct errors through dialogue interaction, thereby ensuring data accuracy and consistency.
[0184] like Figure 20 As shown in FIG, a schematic block diagram of a multimodal scientific research data entry system 2000 according to an embodiment of the present application is shown. Figure 20 As shown, the multimodal scientific research data entry system 2000 of the embodiment of the present application includes an interactive front-end module 2001 and a back-end module 2002.
[0185] The backend module 2002 is used to receive multimodal scientific research data collected by at least one input device through an input interface based on the current scientific research unit plan.
[0186] The multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data and structured data.
[0187] The back-end module 2002 is also used to perform structured processing on the multimodal scientific research data based on data field format specifications, so as to convert the multimodal scientific research data into structured operation instructions; and convert the multimodal scientific research data into structured scientific research records according to the scientific research unit plan.
[0188] The structured operation instruction may be composed of multiple operation instructions, and structured data may be generated when the multiple operation instructions are executed.
[0189] The interactive front-end module 2001 includes a first display area and a second display area, wherein the first display area is used to display and / or play unstructured data, and the second display area is used to display structured data.
[0190] The embodiment of the present application does not require manual filling, and all experimental records are automatically filled in by the AI system, which improves efficiency compared to the traditional manual filling method.
[0191] The following combination Figure 21 The terminal supporting multimodal scientific research data entry of this application is described, wherein: Figure 21 A schematic block diagram of a terminal supporting multimodal scientific research data entry according to an embodiment of the present application is shown.
[0192] like Figure 21 As shown, the terminal 2100 supporting multimodal scientific research data entry includes an input device 2101. The input device 2101 may include at least one of the following: a keyboard, a mouse, a microphone, a scanner, a camera, a handwriting tablet, a drawing tablet, a USB interface, a network card, and the like.
[0193] Continue to combine Figure 21 The terminal 2100 that supports multimodal scientific research data entry may also include: one or more memories 2102 and one or more processors 2103, wherein the memory 2102 stores a computer program run by the processor 2103, and when the computer program is run by the processor 2103, the processor 2103 executes the multimodal scientific research data recording method described above.
[0194] The terminal 2100 supporting multimodal scientific research data entry may be part or all of a computer device that can implement a multimodal scientific research data recording method through software, hardware, or a combination of software and hardware.
[0195] like Figure 21As shown, the terminal 2100 supporting multimodal scientific research data entry includes one or more memories 2102, one or more processors 2103, a display (not shown) and a communication interface, etc. These components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). It should be noted that Figure 21 The components and structure of the terminal 2100 supporting multimodal scientific research data entry shown are merely exemplary and non-restrictive. The terminal 2100 supporting multimodal scientific research data entry may also have other components and structures as needed.
[0196] The memory 2102 is used to store various data and executable program instructions generated during the execution of the relevant methods, such as for storing various application programs or algorithms for implementing various specific functions. It can include one or more computer program products, which can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory. The non-volatile memory can include, for example, read-only memory (ROM), a hard disk, flash memory, etc.
[0197] The processor 2103 can be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can be other components in the terminal 2100 that supports multimodal scientific research data entry to perform desired functions.
[0198] In one example, the terminal 2100 supporting multimodal scientific research data entry also includes an output device that can output various information (such as images or sounds) to the outside (such as an operator), and can include one or more of a display device, a speaker, etc.
[0199] The communication interface can be an interface of any currently known communication protocol, such as a wired interface or a wireless interface, wherein the communication interface may include one or more serial ports, USB interfaces, Ethernet ports, WiFi, wired networks, DVI interfaces, device integrated interconnection modules or other suitable ports, interfaces, or connections.
[0200] In addition, according to an embodiment of the present application, a storage medium is also provided, on which program instructions are stored, and when the program instructions are executed by a computer or processor, the corresponding steps of the multimodal scientific research data recording method of the embodiment of the present application are executed. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.
[0201] In addition, an embodiment of the present application also provides a computer program product, which implements the steps of the above method when the computer program / instructions are executed by a processor.
[0202] The multimodal scientific research data entry system, terminal, storage medium and computer program product supporting multimodal scientific research data entry of the embodiments of the present application have the same advantages as the aforementioned multimodal scientific research data recording method because they can implement the aforementioned multimodal scientific research data recording method.
[0203] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely illustrative and are not intended to limit the scope of the present application. Various changes and modifications may be made therein by those skilled in the art without departing from the scope and spirit of the present application. All such changes and modifications are intended to be included within the scope of the present application as required by the appended claims.
[0204] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0205] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units described is merely a logical function division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another device, or ignoring or not performing some features.
[0206] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and techniques are not shown in detail so as not to obscure the understanding of this description.
[0207] Similarly, it should be understood that in order to streamline the present application and aid in understanding one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, this approach of the present application should not be interpreted as reflecting the intention that the application claimed for protection requires more features than those explicitly recited in each claim. More precisely, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with fewer features than all the features of a single disclosed embodiment. Therefore, the claims following the detailed description are hereby expressly incorporated into the detailed description, with each claim itself serving as a separate embodiment of the present application.
[0208] It will be understood by those skilled in the art that, except where mutually exclusive, all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or apparatus disclosed herein may be combined in any combination. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature providing the same, equivalent, or similar purpose.
[0209] Furthermore, those skilled in the art will appreciate that although some embodiments described herein include certain features included in other embodiments but not other features, combinations of features from different embodiments are intended to be within the scope of this application and to form different embodiments. For example, in the claims, any of the claimed embodiments may be used in any combination.
[0210] The various component embodiments of the present application can be implemented in hardware, or in a software module running on one or more processors, or in a combination thereof. Those skilled in the art will appreciate that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all of the functions of some modules according to the embodiments of the present application. The application can also be implemented as a part or all of a device program (e.g., a computer program and a computer program product) for performing the method described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0211] It should be noted that the above embodiments illustrate rather than limit the present application, and that a person skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbols placed between brackets should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of appropriately programmed computers. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
[0212] The above description is merely a specific embodiment or illustration of a specific embodiment of the present application, and the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. The scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A multimodal scientific research data recording method, characterized in that: Applied to a multimodal scientific research data entry system, the multimodal scientific research data entry system includes a backend module; the method includes: Based on the current scientific research unit scheme, receiving multimodal scientific research data collected by an input device through an input interface; wherein the multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data, and structured data; Based on the data field format specification, calling the back-end module to perform structured processing on the multimodal scientific research data, so as to convert the multimodal scientific research data into structured operation instructions; Converting the structured operation instructions into structured scientific research records according to the scientific research unit plan; Converting the structured operation instruction into a code marked with the structured operation instruction through the back-end module; Executing the code marked with the structured operation instruction to form a structured scientific research experiment record; Obtaining instructions for recording the experimental records in a predetermined country language; Translating the structured operating instructions into the predetermined national language and recording them in the experimental record; Based on the data field format specification, calling the back-end module to perform structured processing on the multimodal scientific research data so as to convert the multimodal scientific research data into structured operation instructions, including: Identifying and classifying the multimodal scientific research data using a data modality identification module to determine the type of the multimodal scientific research data; In the case where the multimodal scientific research data is unstructured data, based on the data type of the multimodal scientific research data, the encoder corresponding to the data type of the multimodal scientific research data in the back-end module is called to convert the multimodal scientific research data into the structured operation instructions.
2. The method according to claim 1, characterized in that The multimodal scientific research data is voice data; the method further comprises: Invoking a speech encoder corresponding to the speech data so that the speech encoder converts the speech data into unstructured text data; The unstructured text data is output.
3. The method according to claim 1, characterized in that The multimodal scientific research data is image data; the method further comprises: Invoking an image encoder corresponding to the image data, so that the image encoder converts the image data into unstructured text data by using a text recognition algorithm and / or an optical character recognition algorithm based on the annotated industry knowledge dataset; The unstructured text data is output.
4. The method according to claim 1, wherein in, The text data is unstructured text data.
5. The method according to any one of claims 2 to 4, characterized in that The multimodal scientific research data entry system further includes an interactive front-end module; the interactive front-end module includes a first display area and a second display area; the method further includes: Sending the unstructured text data to the first display area so that the first display area displays the unstructured text data in the form of a text block; and / or Receive an unstructured text instruction input by an operator in the first display area, and respond to the unstructured text instruction.
6. The method according to claim 5, characterized in that The multimodal scientific research data is structured data; the method includes: sending the structured text data to the second display area so that the second display area displays the structured text data; Among them, the data type and value of the structured data should follow the customized rules of the scientific research unit plan data field input box.
7. The method according to claim 5, wherein The method further comprises: Playing the unstructured text data in a voice interactive manner; and / or Receive the operator's voice command and respond to the voice command.
8. The method according to claim 5, characterized in that Convert the unstructured text instructions into structured operation instructions according to the scientific research unit plan, including: Based on the current scientific research unit program prompt words, a preset algorithm is used to analyze and understand the unstructured text instructions to convert the unstructured text data into corresponding structured operation instructions.
9. The method according to claim 8, characterized in that Based on the current scientific research unit program prompt word, and using a preset algorithm to analyze and understand the unstructured text instruction to convert the unstructured text instruction into a corresponding structured operation instruction, further comprising: An embedded tool specific to the scientific research unit scheme prompt word architecture is used to vectorize each knowledge point in each text block and store it in the form of key-value pairs in the scientific research unit scheme data field input box for subsequent fast matching indexing.
10. The method according to claim 9, characterized in that in, The scientific research unit plan prompt words are pre-set in the back-end module; and / or, the scientific research unit plan prompt words are automatically created by the back-end module based on the knowledge base corresponding to the current scientific research unit plan.
11. The method according to claim 6, characterized in that Convert the structured operation instructions into structured scientific research records according to the scientific research unit plan, including: The structured data is stored in the scientific research unit plan data field input box and displayed in the second display area.
12. The method according to claim 11, characterized in that The method further includes: based on the data field format specification, calling the back-end module to perform structured processing on the multimodal scientific research data, so as to convert the multimodal scientific research data into structured operation instructions, including: Calling the backend module and directly converting the multimodal scientific research data into structured data using the built-in algorithm of the backend module; The structured data is stored in the scientific research unit plan data field input box and displayed in the second display area.
13. The method according to claim 1, wherein The scientific research unit plan data field input box is also used to store multiple data records; the method further includes: Based on the constraint relationship and / or data verification relationship between the multiple data records, the multiple data records are verified, and if there is a logical error in the multiple data records, a prompt indicating the existence of the logical error is sent.
14. The method according to claim 1, wherein The method further comprises: Based on the storage instruction input by the operator, after the experimental record is stored in the data according to the scientific research unit plan, the text record and the text summary stored in the scientific research unit plan data field input box are released: or, Based on the non-storage instruction input by the operator, the text record and the text summary stored in the scientific research unit program data field input box are directly released.
15. The method according to claim 14, characterized in that The method further comprises: Based on the management instructions input by the operator, the content stored in the data field input box of the scientific research unit plan is managed; The management operation includes at least one of the following operations: deleting data, storing data, and modifying data.
16. The method according to claim 5, characterized in that The first display area and the second display area are respectively provided with scroll bars. When the first display area and / or the second display area displays a lot of content, the scroll bars automatically scroll to display the latest content.
17. The method according to claim 1, wherein The input device includes at least one of the following: a keyboard, a mouse, a microphone, a scanner, and a camera.
18. A multimodal scientific research data entry system, characterized in that: The system includes a back-end module and an interactive front-end module; wherein, The back-end module is configured to receive multimodal scientific research data collected by at least one input device through an input interface based on the current scientific research unit plan; perform structured processing on the multimodal scientific research data based on data field format specifications to convert the multimodal scientific research data into structured operation instructions; and convert the structured operation instructions into structured scientific research records in accordance with the scientific research unit plan; Converting the structured operation instruction into a code marked with the structured operation instruction through the back-end module; Executing the code marked with the structured operation instruction to form a structured scientific research experiment record; Obtaining instructions for recording the experimental records in a predetermined country language; Translating the structured operating instructions into the predetermined national language and recording them in the experimental record; The multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data and structured data; The interactive front-end module includes a first display area and a second display area, the first display area is used to display and / or play unstructured data, and the second display area is used to display structured data; Based on the data field format specification, calling the back-end module to perform structured processing on the multimodal scientific research data so as to convert the multimodal scientific research data into structured operation instructions, including: Identifying and classifying the multimodal scientific research data using a data modality identification module to determine the type of the multimodal scientific research data; In the case where the multimodal scientific research data is unstructured data, based on the data type of the multimodal scientific research data, the encoder corresponding to the data type of the multimodal scientific research data in the back-end module is called to convert the multimodal scientific research data into the structured operation instructions.
19. A terminal supporting multimodal scientific research data entry, characterized in that: The terminal includes: At least one input device, wherein the input device collects multimodal scientific research data and sends the multimodal scientific research data to the processor for processing; the input device includes at least one of the following: a keyboard, a mouse, a microphone, a scanner, and a camera; A memory and the processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the processor executes the multimodal scientific research data recording method according to any one of claims 1 to 17.
20. A storage medium, characterized in that The storage medium stores a computer program, which, when executed by a processor, enables the processor to execute the multimodal scientific research data recording method according to any one of claims 1 to 17.
21. A computer program product, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 17 are implemented.