Multi-mode scientific research data recording method and system, terminal and storage medium

By structured processing of multimodal scientific research data and converting it into structured operation instructions, the problems of inefficiency of traditional experimental recording methods and intimate data integration are solved, and efficient and accurate experimental recording is achieved.

CN120011362AActive Publication Date: 2025-05-16WESTLAKE UNIV

Patent Information

Application Number
CN202510086467.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-16
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The traditional experimental recording method is inefficient, complex in operation, and the correlation and integration between different modal data are not tight enough, which affects the comprehensiveness and accuracy of experimental recording.

Method used

It provides a multimodal scientific research data recording method, which receives multimodal scientific research data through the input interface, calls back-end module for structured processing, converts it into structured operation instructions, and converts it into structured scientific research records according to the scientific research unit plan.

Benefits of technology

It reduces manual input time, improves experimental recording efficiency, automatically recognizes and records multimodal scientific research data, and corrects errors through dialogue and interaction to ensure data accuracy and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011362A_ABST
    Figure CN120011362A_ABST
Patent Text Reader

Abstract

The invention provides a multi-mode scientific research data recording method and system, a terminal and a storage medium. The method is applied to a multi-modal scientific research data entry system, and the multi-modal scientific research data entry system comprises a rear-end module. The method comprises the following steps: based on a current scientific research unit scheme, receiving multi-modal scientific research data collected by an input device through an input interface; the multi-modal scientific research data comprises at least one of the following data forms: text data, voice data, image data and structured data; based on a data field format specification, calling the back-end module to perform structured processing on the multi-modal scientific research data so as to generate a structured operation instruction; and converting the structured operation instruction into a structured scientific research record according to a scientific research unit scheme. According to the embodiment of the invention, manual input time can be reduced, experimental recording efficiency can be improved, multi-modal scientific research data can be automatically identified and recorded, errors are corrected through dialogue interaction, and data accuracy and consistency are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data analysis technology, and in particular to a multimodal scientific research data recording method, system, terminal and storage medium. Background Art

[0002] Although traditional experimental recording methods can import text data, voice data, handwritten records taken from photos, and standardized data output by experimental instruments and equipment, etc., these data still need to be manually input by researchers or imported into the recording platform in a relatively passive way, which leads to insufficient recording efficiency and convenience, and problems such as low recording efficiency and complex operation. Especially when the experimental process is busy or needs to be recorded quickly, the convenience of existing tools cannot meet the needs.

[0003] Moreover, although it can support multiple modal data, the association and integration between different modal data are not tight enough, which affects the comprehensiveness and accuracy of experimental records. In addition, traditional experimental records are strictly written in the format of templates and are not flexible enough. Summary of the invention

[0004] In view of at least one of the above technical problems existing in the prior art, this application is proposed. According to the first solution of this application, a multimodal scientific research data recording method is provided, which is applied to a multimodal scientific research data entry system, and the multimodal scientific research data entry system includes a backend module; the method includes:

[0005] Based on the current scientific research unit scheme, receiving multimodal scientific research data collected by an input device through an input interface; wherein the multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data and structured data;

[0006] Based on the data field format specification, calling the back-end module to perform structured processing on the multimodal scientific research data, so as to convert the multimodal scientific research data into structured operation instructions;

[0007] The structured operation instructions are converted into structured scientific research records according to the scientific research unit plan.

[0008] The multimodal scientific research data recording method of the embodiment of the present application structures the multimodal scientific research data to form structured operation instructions, and converts the structured operation instructions into structured scientific research records according to the scientific research unit plan. This can reduce manual input time and improve experimental recording efficiency, and automatically identify and record multimodal scientific research data, and correct errors through dialogue interaction, thereby ensuring data accuracy and consistency.

[0009] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below.

[0010] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention, as claimed. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0012] Figure 1 A schematic flow chart showing a multimodal scientific research data recording method according to an embodiment of the present application;

[0013] Figure 2 A schematic flowchart of step S102 according to an embodiment of the present application is shown;

[0014] Figure 3 A schematic diagram showing a multimodal scientific research data recording method using an encoder to pre-process multimodal scientific research data according to an embodiment of the present application;

[0015] Figure 4 A schematic flowchart of step S102 according to another embodiment of the present application is shown;

[0016] Figure 5 A schematic diagram showing a method for recording multimodal scientific research data without using an encoder to pre-process the multimodal scientific research data according to another embodiment of the present application;

[0017] Figure 6 A schematic flow chart showing the use of the first display area to display unstructured text data according to an embodiment of the present application is shown;

[0018] Figure 7 A schematic diagram showing playing unstructured text data in a voice interactive manner according to an embodiment of the present application;

[0019] Figure 8 A schematic flowchart of step S102 according to an embodiment of the present application is shown;

[0020] Fig. 9 A schematic flowchart of step S801 according to an embodiment of the present application is shown;

[0021] Fig.10 A schematic diagram showing an operation instruction for converting text data into structured data fields based on a scientific research unit scheme prompt word according to an embodiment of the present application;

[0022] Fig.11 A schematic diagram showing a method of converting picture data into structured data field operation instructions using a picture encoder based on a scientific research unit solution prompt word according to an embodiment of the present application;

[0023] Fig.12 A schematic diagram showing a display interface according to an embodiment of the present application;

[0024] Fig.13 A schematic flowchart of step S103 according to another embodiment of the present application is shown;

[0025] Fig.14 A schematic flow chart showing the conversion of a structured operation instruction into a code marked with a structured operation instruction according to an embodiment of the present application is shown;

[0026] Fig.15 A schematic flow chart showing verification of data records based on data verification relationships and / or data constraint relationships between / among multiple data fields according to an embodiment of the present application is shown;

[0027] Fig.16 A schematic flow chart showing permanent storage and release of data records according to an embodiment of the present application is shown;

[0028] Fig.17 A schematic flow chart showing the management operation of the content of the scientific research unit program data field input box according to an embodiment of the present application;

[0029] Fig.18 A schematic flow chart showing recording an experiment in a predetermined national language according to an embodiment of the present application;

[0030] Fig.19 A schematic diagram showing the interaction between an operator, an AI interaction front end, a data recording end, an AI back end and an AI database according to an embodiment of the present application is shown;

[0031] Fig. 20 A schematic block diagram of a multi-modal scientific research data entry system according to an embodiment of the present application is shown;

[0032] Fig.21 A schematic block diagram of a terminal supporting multimodal scientific research data entry according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0033] In order to enable those skilled in the art to better understand the technical solutions of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0034] Experimental records in the field of natural sciences are usually composed of data from different sources, different modalities, and different representations (for example, handwritten records, spreadsheets, instrument photos, etc.). At the same time, in some special environments and scenarios, researchers need to be able to use free forms such as voice interaction and dialogue in non-restricted scenarios to complete experimental records. However, it is difficult to share and reuse data in different modalities.

[0035] Based on at least one of the aforementioned technical problems, the present application provides a multimodal scientific research data recording method, which is applied to a multimodal scientific research data entry system, wherein the multimodal scientific research data entry system includes an interactive front-end module and a back-end module: the interactive front-end module is mainly used for human-computer interaction; the back-end module is mainly used for structured processing and storage of multimodal scientific research data. The method includes: based on the current scientific research unit scheme, receiving multimodal scientific research data collected by an input device through an input interface of the interactive front-end module; wherein the multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data and structured data; based on the scientific research unit scheme and data field format specifications, calling the back-end module to perform structured processing on the multimodal scientific research data so as to convert the multimodal scientific research data into structured operation instructions; converting the structured operation instructions into structured scientific research records according to the scientific research unit scheme. The multimodal scientific research data recording method of the embodiment of the present application structures the multimodal scientific research data to form structured operation instructions, and converts the structured operation instructions into structured scientific research records according to the scientific research unit plan. This can reduce manual input time and improve experimental recording efficiency, and automatically identify and record multimodal scientific research data, and correct errors through dialogue interaction, thereby ensuring data accuracy and consistency.

[0036] The execution subject of the embodiment of the present application may be a multimodal scientific research data entry system. The multimodal scientific research data entry system may be an artificial intelligence (AI) system, and the AI ​​system may be based on large language models (LLMs) or large multimodal models (LMMs), and this article does not distinguish between the three. The AI ​​system can realize structured processing of multimodal scientific research data; it can not only automatically record experimental data, but also interact with operators, correct various numerical errors and logical errors, etc., to ensure data accuracy and consistency.

[0037] The large language model in this article refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. The large language model can handle a variety of natural language tasks, such as text classification, question and answer, dialogue, etc. At present, the large language model adopts a transformer (Transformer) architecture and pre-training objectives similar to the small model. The difference from the small model is the increase in model size, training data and computing resources. The large language model (LLM) of the embodiment of the present application may further include a variety of encoders corresponding to a variety of modalities, and each encoder can convert the data of the corresponding modality into text or its equivalent representation (such as token or vector embedding), so that the large language model has the ability to understand and process multimodal information. The large language model can thus process multimodal scientific research data and generate a series of executable structured operation instructions from it.

[0038] The multimodal big model in this article refers to the combination of big models with multimodal data such as text, images, audio, and video for training, so as to achieve the fusion and understanding of multimodal information in a unified representation space; with this fusion capability, the multimodal big model can not only process the semantics at the text level, but also conduct more in-depth analysis and reasoning on non-textual data such as images, audio, and video, and then demonstrate more comprehensive intelligent processing capabilities in complex tasks such as cross-modal retrieval, multimodal question and answer, and image and text generation. Under the paradigm of the multimodal big model, the information contained in multiple modal inputs can be directly handed over to the multimodal big model for processing, and it generates a series of executable structured operation instructions.

[0039] Figure 1 A schematic flow chart of a multimodal scientific research data recording method according to an embodiment of the present application is shown; Figure 1 As shown, the multimodal scientific research data recording method 100 according to the embodiment of the present application may include the following steps S101, S102 and S103:

[0040] In step S101, based on the current research unit protocol, multimodal research data collected by an input device is received through an interactive front-end module input interface.

[0041] The multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data and structured data. The multimodal scientific research data of the embodiment of the present application can provide operators with a more intuitive and detailed description of research activities.

[0042] The text data in multimodal scientific research data can be text data input by the operator through the input interface in the form of human-computer dialogue with the large language model. This text data can be text expressed in natural language. Voice data can be a voice input by the operator through the input interface, or it can be in the form of audio or recording; image data can be an image collected from the experimental object, or it can be a picture taken from a piece of text written by humans, or it can be an image of the equipment, etc., or it can be a video data. Structured data can be structured data generated by experimental instruments / equipment, or structured data derived from instruments / equipment. Structured data usually has a clear predefined data model and structure, can be stored in a relational database, and can be queried and operated using SQL (Structured Query Language). The characteristics of structured data are highly organized, data is stored in the form of rows and columns, has a fixed format and length, and each field has a predefined data type, such as integer, string, date, etc. The advantages of structured data include easy retrieval, update and deletion operations, as well as data consistency and accuracy. Common structured data include values, dates, times, telephone numbers, addresses, etc. The structured data may also include structured data generated by an instrument application programming interface (API) / computational model, etc. In the embodiment of the present application, the operator may directly input the structured data through the interactive front-end module.

[0043] Among them, the input interface can be a series of human-machine interfaces defined by an operator, and multimodal scientific research data can be received through the input interface.

[0044] The input device includes at least one of the following: a keyboard, a mouse, a microphone, a scanner, a camera, a handwriting tablet, a drawing board, a USB interface, a network card, etc.

[0045] The essence of the scientific research unit scheme here is composed of components such as scientific research protocol (Protocol), model (Model), assigner (Assigner), etc. In the scientific research unit scheme Protocol, operators can customize various types of data fields and their corresponding data field format specifications (Data Field JSON Schema).

[0046] In step S102, based on the scientific research unit scheme and data field format specification, the back-end module is called to process the multimodal scientific research data so as to convert the multimodal scientific research data into structured operation instructions.

[0047] The structured operation instructions here refer to a series of executable single / multiple operation instructions.

[0048] In an embodiment of the present application, the multimodal scientific research data involved in the current scientific research unit scheme can be sent to the back-end module for processing. The large language model in the embodiment of the present application may include a speech encoder, a picture encoder, a structured data encoder, and other encoders. The data of each modality corresponds to its own encoder. For example, speech data corresponds to a speech encoder, and image data corresponds to a picture encoder.

[0049] It is worth noting that after parsing the various data, the above-mentioned encoders will convert the various data into text data according to the scientific research unit scheme. Then the large language model processes the text data to obtain structured operation instructions. At the same time, since the above-mentioned encoders parse the various data into text data, for the text data in the multimodal scientific research data, there is no need to provide additional encoders for parsing operations, but the text data can be sent directly to the large language model for processing.

[0050] In one embodiment of the present application, Figure 2 As shown, step S102 includes step S201, step S202 and step S203:

[0051] In step S201, the multimodal scientific research data is identified and classified by a data modality identification module to determine the type of the multimodal scientific research data;

[0052] In step S202, based on the data type of the multimodal scientific research data, an encoder and a large language model corresponding to the data type of the multimodal scientific research data in a backend module are called to convert the multimodal scientific research data into the structured operation instruction;

[0053] In step S203, the backend processing module converts the structured operation instructions into structured scientific research data.

[0054] The structured operation instructions of the embodiment of the present application can be composed of multiple operation instructions that can be directly executed, and when the multiple operation instructions are executed, structured data can be generated. The structured data will be stored in the data field input box of each scientific research unit plan and finally form an experimental record.

[0055] For example, the structured operation instructions expressed in a certain programming language are as follows:

[0056]

[0057] Since the structured research data in a research record can actually be stored as a series of key-value pairs under the research unit scheme paradigm, the structured research data generated based on the above operation instructions is:

[0058]

[0059] That is to say, the experimenter is updated to "Zhang San", the experimental temperature is updated to 25 degrees, and the experimental humidity is updated to 50%.

[0060] The core technical solutions of the embodiments of the present application are to convert the data in the multimodal scientific research data input such as text / voice / picture into structured operation instructions, so that the system can insert the relevant information of the structured scientific research data contained therein into the corresponding data fields based on these structured operation instructions.

[0061] In one embodiment of the present application, the multimodal scientific research data may be structured according to the data type of the multimodal scientific research data. Figure 3 FIG. 1 is a schematic diagram of processing multimodal scientific research data using a corresponding type of encoder based on the type of multimodal scientific research data according to an embodiment of the present application. The following is an example of the processing process of different types of multimodal scientific research data.

[0062] In the first example, the multimodal scientific research data is speech data. First, a speech encoder corresponding to the speech data is called so that the speech encoder converts the speech data into text. Then, the text is output. For example, in combination with Fig.12 When the operator speaks the voice of "search PCR", the voice encoder converts the voice data into the text form of "search PCR" and presents it in the first display area.

[0063] In the second example, the multimodal scientific research data is image data. First, an image encoder corresponding to the image data is called, so that the image encoder converts the image data into text based on the annotated industry knowledge data set through a text recognition algorithm and / or an optical character recognition algorithm. Then, the text is output.

[0064] Continue to combine Figure 3 For example, image data may include handwritten text photos and equipment photos. For handwritten text photos, a text recognition algorithm (textract algorithm) may be used to extract text from handwritten text photos and recognize the handwritten text as text. For experimental objects, a text recognition algorithm (textract algorithm) and an optical character recognition algorithm (OCR technology) may be used to perform image analysis on equipment photos, extract text and target objects therein, and then generate text. For equipment photos, an optical character recognition algorithm (OCR technology) may be used to perform image analysis on equipment photos, obtain the state of the equipment, and then generate text. For example, after image analysis, it is found that the lid of a certain device is not closed properly, and "the lid of a certain device is not closed properly" may be recorded in the experimental record in the form of text. This method can not only directly convert image data into text in the terminal, but also prompt for irregular behavior of experimental operations to improve the accuracy of the experiment.

[0065] Combination Fig.12 For example, before the formal experiment, the operator needs to record the experimental environment. For example, a certain experimental equipment is photographed and then input into the AI ​​system. The backend of the AI ​​system analyzes the photo and finds that the cover of a certain experimental equipment is not closed. The AI ​​system will record this information and display it in the second display area 1202 in the form of a text annotation, that is, "the cover of a certain equipment is not closed."

[0066] It is worth noting that video data is composed of multiple frames of pictures. Video data belongs to image data, and the processing of video data is consistent with the processing of image data.

[0067] Among them, image data is an important unstructured data source. In scientific experiments, it is often necessary to take pictures of the experimental process, and operators sometimes need to handwrite experimental records, etc. The embodiment of the present application will call a large language model to identify image-based scientific research data and enter it into the scientific research unit program data field input box (RU Data Field InputBox) for storage. The processing of image data is mainly divided into two situations: first, in the initialization stage, the image information is used for initialization to identify and record the scientific research data contained in the image information; second, consistent with the above-mentioned scenario in the experimental record of text data, the input of image data is supported in the dialogue.

[0068] In the fourth example, the multimodal scientific research data is structured data. The present application embodiment also supports experimental records of structured markup documents such as Word, JSON, Markdown, etc., and these structured documents can be identified and entered using large language models and structured parsing tools.

[0069] In an embodiment of the present application, the operator can manually fill in the structured data in the scientific research unit program data field input box according to the current scientific research unit program. For example, the operator manually fills in the experimental data such as the experimenter, experiment name, experiment date, experiment temperature, experimental object, and change value of the experimental object in the scientific research unit program data field input box. It is worth noting that when manually filling in the structured data, it can be directly filled in the scientific research unit program data field input box in the second display area so that the second display area displays the structured data. Among them, the data type and value of the structured data should follow the rules customized by the scientific research unit program data field input box.

[0070] Regarding the experimental records of instrument equipment and computing model outputs, since the instruments and computing models widely used in modern scientific experiments can automatically generate a large amount of data, such as experimental measurement values, instrument status, statistical results, etc., the embodiments of the present application can enable the large language model to connect to the APIs of such instruments, equipment and models to obtain input and output information and perform experimental records.

[0071] In addition, the structured operation instructions generated by the large language model are texts. When data such as voice data, image data and structured data are obtained, these data need to be converted into text (such as voice, image, video and other data encoded in Base64 string format; or voice, image, video and other data are uploaded to the local database / cloud database / local area network / Internet to obtain the corresponding text form ID / URL / URI), and then the text is stored in the scientific research unit program data field input box, and the text is displayed on the display interface. If the structured operation instructions generated by the large language model contain data in the form of voice, these data in the form of voice can be stored in the recording slot, and when the play command is received, they can be played. If the structured operation instructions generated by the large language model contain data in the form of pictures, these data in the form of pictures can be stored in the recording slot, and when the display command is received, they can be displayed in the display interface or preview window. If the structured operation instructions generated by the large language model contain data in the form of videos, these data in the form of videos can be stored in the recording slot, and when the play command is received, they can be played through the video player. If the structured operation instructions generated by the large language model contain data in the form of structured documents, these data in the form of structured documents can be stored in the record slot, and when receiving the parsing or display instructions, they can be automatically parsed and presented according to the preset format requirements.

[0072] The embodiments of the present application can realize intelligent experiment recording on a general extensible platform, that is, automatically identify and record standardized data such as text data, voice data (including voice commands), handwritten text pictures, and output of experimental instruments and equipment.

[0073] In another embodiment of the present application, Figure 4 As shown, step S102, based on the data field format specification, calls the back-end module (including the multimodal large model) to perform structured processing on the multimodal scientific research data, so that the multimodal scientific research data is converted into structured operation instructions (single / multiple operation instructions), including step S401 and step S402:

[0074] In step S401, the backend module is called, and the multimodal scientific research data is directly converted into structured operation instructions using the built-in algorithm of the backend module;

[0075] In step S402, the structured scientific research data contained in the structured operation instruction is stored in the scientific research unit program data field input box and displayed in the second display area.

[0076] like Figure 5 As shown, after receiving the multimodal scientific research data (text data, voice data, image data, instrument API / computational model and other data), the input device sends the multimodal scientific research data directly to the multimodal large model, so that the multimodal large model uses the built-in algorithm for processing. For example, the multimodal large model can use a deep learning algorithm to analyze and understand the multimodal scientific research data, generate multiple structured operation instructions, and automatically execute multiple structured operation instructions to generate corresponding structured text data. These structured text data will be stored in the scientific research unit program data field input box. In addition, after analyzing and understanding the multimodal scientific research data, the multimodal large model can output unstructured text data and display the structured text data on the display interface (such as the first display area) so that the user can intuitively understand the content of the experimental record.

[0077] In one example, combining Fig.12 , Fig.12 The second display area 1202 in the display is the structured text data obtained by executing the structured operation instruction, and these structured text data will be presented in the final experiment record. The user can directly enter the corresponding content, i.e., structured text data, in the scientific research unit program data field input box corresponding to data fields such as "experimenter", "experiment date", "experiment purpose", "temperature", and "humidity".

[0078] In another example, continue with Fig.12 , Fig.12The first display area 1201 in the figure displays unstructured text data, which are obtained by the AI ​​system after analyzing and understanding multimodal data (such as voice, pictures, text, etc.). These unstructured text data need to be processed again before they can be presented as structured text data displayed in the second display area 1202. For example, the operator asks the question "Please change the temperature to 45 degrees" in the form of a voice question. After analyzing and understanding, and executing the operation, the AI ​​system answers in the form of a voice "The temperature has been changed to 45 degrees".

[0079] That is, in this example, the AI ​​system converts the multimodal data into unstructured text data (e.g., answers generated by the AI ​​based on user questions) and presents it in the first display area 1201. It is worth noting that the user can enter questions in the form of text data. For example, if the user enters the question "Search for PCR" in the question input box of the first display area 1201, the AI ​​system will not be able to find PCR after searching, and will display "Unable to find PCR" as text data.

[0080] It is worth noting that the current scientific research unit solution not only performs intuitive language conversion for the input multimodal scientific research data, but also can understand the input data based on algorithms such as deep learning algorithms. Any problems that arise in the actual process can be recorded and displayed in the final experimental record. For example, when a photo (or video) of a certain device is parsed and it is found that the cover of a certain device is not properly closed, it will be recorded in the data field input box of the scientific research unit solution, "The cover of a certain device is not properly closed". The embodiment of the present application can parse the photo, obtain the state of the object in the photo, and record the state of the object in the data field input box of the scientific research unit solution to facilitate the operator to make timely adjustments.

[0081] The AI ​​system of the embodiment of the present application can better integrate, analyze and understand multimodal data. For example, for the image and voice data obtained during the experiment through the camera and microphone, the image encoder and voice encoder can be used to convert them into text and then use the large language model for analysis and understanding, or directly use the multimodal large model for analysis and understanding, and generate corresponding experimental records. In addition, because the understanding of the large language model / multimodal large model has high accuracy and flexibility, it can answer questions related to the experiment and assist the operator in recording the experiment.

[0082] like Figure 6 As shown, the method further includes step S601 and step S602:

[0083] In step S601, the unstructured text data is sent to the first display area, so that the first display area displays the unstructured text data in the form of a text block; and / or

[0084] In step S602, an unstructured text instruction input by an operator in the first display area is received, and a response is made to the unstructured text instruction.

[0085] In one embodiment of the present application, the multimodal scientific research data entry system further includes an interactive front-end module. Fig.12 As shown, the display interface 1200 of the interactive front-end module includes a first display area and a second display area. The first display area is used to display the human-computer question and answer content in text form, or to display voice data (which has been parsed into text data by the voice encoder), image data (which has been parsed into text data by the image encoder), and structured data (which has been parsed into text data by the structure encoder). Please refer to the following for a detailed introduction to the first display area and the second display area.

[0086] For example, the voice interaction data is converted into text using natural language processing (NLP) technology to generate text corresponding to the voice; the voice interaction data may include voice instructions, requiring the AI ​​system to record scientific research data, and / or answer questions related to the scientific research unit plan and record key information during the experiment.

[0087] In some embodiments, the content displayed in the first display area can be played out in the form of voice while being displayed, so that the operator can know the displayed content and the feedback of the AI ​​system without having to go to the display interface. In addition, the user can also directly use voice to give command feedback to the broadcast voice, so that the AI ​​system can further process and give feedback based on the user's further voice command. Figure 7 As shown, the method further includes step S701 and step S702:

[0088] In step S701, the unstructured text data is played in a voice interactive manner; and / or

[0089] In step S702, a voice instruction from an operator is received and a response is given to the voice instruction.

[0090] During multiple rounds of human-computer question-answering (voice interaction), the large language model can use natural language processing technology to perform speech-to-text conversion, text summary generation and other operations on the voice interaction data. The operator can input the experimental records through voice, and the large language model will convert the voice into text and store it. Through the NLP-based dialogue system, it can interact with the operator, answer questions related to the experiment, and assist in recording key information during the experiment.

[0091] In some embodiments, the operator can also issue some operation instructions to the large language model. After executing these operation instructions, the AI ​​system can provide feedback to the operator on the execution status. For example, when the operator starts the experiment, he says "please initialize", and the conversation module (for example, the dialogue robot) in the AI ​​system obtains the initialization status of each experimental instrument or equipment based on the generated text summary "initialize". When all experimental instruments or equipment have completed the initialization operation, the voice answers "initialization is complete". For another example, the operator says "such and such experiment, Zhang San", the AI ​​system can understand "Zhang San" as the experimental operator or experimental recorder according to the natural language processing algorithm, and then record "Operator: Zhang San" or "Experimental recorder: Zhang San" in the experimental record.

[0092] In addition, based on the operator's customized rules, the AI ​​system can be configured to maintain the human-computer question and answer process throughout the entire experimental process, so as to record the data generated by multiple experimental steps in a single experiment or multiple scientific research data from multiple experiments.

[0093] In addition, the AI ​​system can not only answer the operator's questions, but also search based on the keywords of the questions. In the embodiment of the present application, the operator can issue instructions to the AI ​​system in the form of multiple rounds of dialogue to perform operations such as modifying data, correcting data, deleting data, storing data, and updating data on the experimental process data based on the experimental process data displayed on the display interface.

[0094] For example, when the operator says "help me search for PCR", "PCR" is searched as the field name. If "PCR" is not found, the operator can provide text / voice feedback "PCR not found, please try again". The operator can set the search times to 3 times before answering or stop searching if no results are found after 3 searches. The embodiment of the present application can perform semantic analysis on voice data, execute instructions according to their meaning, and then convert the execution results into text and provide feedback to the operator in the form of voice data.

[0095] In addition, the operator can conduct multiple rounds of questions and answers with the AI ​​system at any time during the entire experiment, and the question-and-answer process and results can be stored and displayed in the scientific research unit program data field input box. In view of the problem that operators cannot free their hands to record scientific research data in certain specific scenarios, the embodiment of the present application can also record the experimental process based on voice interaction, using the AI ​​system to convert the operator's voice input and model output into the above-mentioned scenes in the experimental record of text data to complete the experimental record based on voice dialogue instructions.

[0096] For example, the user can give the system a command through voice: "Please start recording. Record the experimenter as Zhang San". The system replies: "The experimenter has been recorded as Zhang San", and reports it in voice form. After listening to the voice, the user can continue to give commands through voice: "Please record the experiment date as January 1, 2024". The system immediately replies: "The experiment date has been recorded as January 1, 2024". In this way, the user can directly record scientific research through the voice control system without approaching the recording interactive terminal or even checking the interactive interface. In this way, it can effectively deal with situations in special experimental scenarios where both hands cannot be freed for keyboard operations. For example, when the user is conducting a chemical experiment in a glove box, he needs to continue to operate inside the glove box and cannot use the keyboard to enter data.

[0097] Therefore, the embodiments of the present application can adapt to the diverse needs of scientific research data and effectively improve the convenience and accuracy of scientific researchers' operations during the recording process.

[0098] In step S102, the unstructured text instructions may also be converted into structured operation instructions according to the scientific research unit plan.

[0099] like Figure 8 As shown, step S102 converts the unstructured text instructions into structured operation instructions according to the scientific research unit scheme, including step S801:

[0100] In step S801, based on the current scientific research unit program prompt words, a preset algorithm is used to analyze and understand the unstructured text instructions to convert the unstructured text data into corresponding structured operation instructions.

[0101] For example, a user issues the following unstructured text instruction:

[0102] "Please set the temperature to 25℃."

[0103] Through S801, it can be converted into the following structured operation instructions:

[0104]

[0105] In one embodiment of the present application, the encoder converts the multimodal scientific research data into the text based on the current scientific research unit scheme prompt word.

[0106] The scientific research unit program prompt words here can be automatically generated based on the current scientific research unit program, or preset by the operator or other operators according to the current scientific research unit program. For example, the prompt words "temperature range is 20℃-50℃" can be displayed below the experimental temperature field. When the obtained temperature is not within the temperature range, a temperature error prompt will be sent.

[0107] like Fig. 9 As shown, step S801 is based on the current scientific research unit program prompt word, and uses a preset algorithm to analyze and understand the unstructured text data to convert the unstructured text data into corresponding structured operation instructions, further including step S901:

[0108] In step S901, an embedded tool specific to the scientific research unit scheme prompt word architecture is used to vectorize each knowledge point in each text block, and store it in the form of key-value pairs in the scientific research unit scheme data field input box for subsequent fast matching indexing.

[0109] Wherein, the scientific research unit plan prompt words are pre-set in the back-end module; and / or, the scientific research unit plan prompt words are automatically created by the back-end module according to the knowledge base corresponding to the current scientific research unit plan.

[0110] In one embodiment of the present application, Fig.10 and Fig.11 This article will introduce how to convert the unstructured text data into corresponding structured operation instructions based on the scientific research unit program prompts. Fig.10 A schematic diagram showing the conversion of text data into structured operation instructions based on scientific research unit program prompt words according to an embodiment of the present application is shown; Fig.11 A schematic diagram of converting image data into structured operation instructions according to the scientific research unit scheme prompt words in an embodiment of the present application is shown. The scientific research unit scheme prompt words are created by the AI ​​system according to the current scientific research unit scheme. The AI ​​system can also construct the scientific research unit scheme background prompt words by itself according to the background knowledge of the experiment. Background knowledge may include basic information defined by the operator, working environment, experimental tasks, experimental purposes, etc. For example, before conducting a medical experiment to treat lung disease, basic information such as room temperature and whether working in a sterile environment is input into the AI ​​system, from which the current scientific research unit scheme background prompt words can be extracted and added to the original scientific research unit scheme prompt words.

[0111] like Fig.12As shown, it is a schematic diagram of the display interface 1200 of the multimodal scientific research data recording method of an embodiment of the present application. The display interface 1200 implemented in the present application includes a first display area 1201 and a second display area 1202. The first display area 1201 is located on the left side, and is used to display the content of the human-computer question and answer; the second display area 1202 is located on the right side, and is used to display the content temporarily stored in the scientific research unit program data field input box. In addition, the first display area 1201 and the second display area 1202 are respectively provided with scroll bars to facilitate the user to manually / automatically scroll the scroll bar when the human-computer question and answer content and the scientific research unit program data field input box and its content are large, so as to display more content.

[0112] like Fig.13 As shown, step S103 converts the structured operation instruction into a structured scientific research record according to the scientific research unit scheme, and also includes step S1301:

[0113] In step S1301, the structured data is stored in the scientific research unit program data field input box and displayed in the second display area.

[0114] Combination Fig.12 , since different data types (such as strings, integers, floating point numbers, Boolean values, dates, enumeration values, etc.) are defined for different fields in the data field format specification, in this case, different scientific research unit program data field input boxes can be generated for the data fields in the second display area, and corresponding interactive controls can be generated for the scientific research unit program data field input boxes according to the data types specially annotated for each field in the data field format specification. For example, annotations based on data types can be provided for the scientific research unit program data field input boxes based on the data type rules of the data field format specification. The embodiment of the present application can help operators focus their time and energy on the definition and development of substantive content such as scientific research protocols and data fields in scientific research programs through this display method, without having to worry about issues such as experimental record interfaces, scientific research data storage structures and methods, so that scientists can efficiently design high-quality scientific research programs that meet actual scientific research needs in an operator-friendly manner in daily scientific research activities, and use them for scientific research data recording.

[0115] Various scientific research data are stored in the database in the form of fields. It is worth noting that in the embodiment of the present application, since the AI ​​system can directly obtain the scientific research data output by various experimental instruments or equipment, all the data in the various fields on the right side of the display interface can be directly displayed without manual filling, which improves efficiency compared to the traditional manual filling method. Moreover, the embodiment of the present application displays all the data in a visual form, which is convenient for operators to see the scientific research data at any time. In the case of a large amount of data on the right, the scientific research data can be automatically sorted and the automatic scrolling interface display can be set, which also greatly improves the efficiency of experimental recording.

[0116] As shown in the figure, Fig.14 As shown, the method further includes step S1401 and step S1402:

[0117] In step S1401, the structured operation instruction is converted into a code marked with the structured operation instruction by the back-end module;

[0118] In step S1402, the code marked with the structured operation instruction is executed to form a structured scientific research record.

[0119] The code marked by the structured operation instruction can be represented as a plurality of operation instructions represented in a certain machine language and capable of being automatically executed by the terminal.

[0120] In one embodiment of the present application, the scientific research unit scheme data field input box is also used to store data corresponding to multiple data fields. The data corresponding to multiple data fields refers to data in corresponding data fields extracted from multimodal scientific research data such as text / voice / picture input by the user.

[0121] For example, if the user says the temperature is set to 45 degrees and the humidity is 50%, then the corresponding processed structured scientific research data is

[0122] {

[0123] "temperature":25.0,

[0124] "humidity":50.0

[0125] }

[0126] The 25.0 and 50.0 here are the corresponding data in the data field.

[0127] like Fig.15 As shown, the method further includes step S1501:

[0128] In step S1501, based on the data verification relationship and / or constraint relationship between / among the multiple data fields, the data corresponding to the multiple data fields are verified, and if there is a logical error in the data corresponding to the multiple fields, a prompt of the existence of the logical error is sent.

[0129] Specifically, the type constraint includes constraining the corresponding data field to use predefined multimodal scientific research data during data entry, wherein the predefined multimodal scientific research data includes one or more of text data, image data, video data, audio data, and text data. In some embodiments, the operator can further define the data type (e.g., various numerical types, time types, etc.) of the data field previously defined in the scientific research protocol. For example, the solvent_volume data field is defined as a floating point number. Then, when the operator enters non-floating point data (e.g., a string of letters) in the field, the system will indicate a type error.

[0130] The numerical verification relationship includes constraining the corresponding data field to follow the specified pattern and / or not exceed the preset value range during data entry. The operator can further define validation rules for the data fields defined in the scientific research agreement, for example, further add validation rules for the solvent_volume data field and its type constraint (floating point number) to ensure that the floating point number filled in the data field must be greater than zero. In this case, if the operator enters a negative floating point number, the system will display a numerical verification error. In another example, the so-called following the specified pattern can be, for example, using RegExp regular expressions RE (Regular Expression) to constrain the composition rules and patterns of the strings in the field, for example, constraining a field related to the mailbox to record a value that must contain and only have 1 "@" symbol, and so on. In other examples, other patterns and value range constraints can also be set, which are not listed here one by one.

[0131] The combined verification relationship includes constraining the type and / or value of each data field to satisfy a predetermined constraint relationship. For example, it can be constrained that in an experiment, when the temperature required by the experiment is greater than 40 degrees, the humidity required by the experiment must be lower than 20%. It can be seen that the temperature-related values ​​and the humidity-related values ​​form a combined verification relationship and are interdependent.

[0132] This allows for rapid identification of abnormal scientific research data. If any verification relationship fails, the reason for the verification failure will be displayed, and the operator can correct the value of the invalid data field one by one according to the error prompt until it is modified to a valid record. Through the above verification process, the platform can ensure that the data entered by the operator meets the requirements.

[0133] In some other embodiments, if the operator defines multiple data fields with dependency and assignment relationships in the scientific research protocol, then these relationships can be further customized using an assigner. Specifically, when the operator designs the scientific research unit plan of the corresponding discipline based on the scientific research node design environment, it can also include an assigner for the data field, and the assigner is used to assign values ​​to the data field based on the data field dependency graph and the assignment rules. In some embodiments, the data field dependency graph is a single-level or multi-level directed acyclic graph, and the assignment relationship between the defined data fields is single dependency or multiple dependency, and each data field is assigned by at most one assigner. An assigner can take one or more upstream data fields as dependencies, and at the same time, a data field can actually be used as a dependency of one or more data fields, but for a specific data field, the way to determine its field value should be unique. Just as an example, if two data fields solvent_volume and solvent_volume_2 are defined in the scientific research protocol, and the operator uses the assigner to define that solvent_volume_2 is always twice the value of solvent_volume, in this case, whenever a new value is entered for solvent_volume, the system will automatically set solvent_volume_2 to twice the solvent volume value. For example, if the value of solvent_volume is 5, the system will automatically assign a value of 10 to solvent_volume_2. This can greatly improve the efficiency and accuracy of scientific research data recording.

[0134] By defining the model of the scientific research unit plan, the operator's input can be dynamically verified to ensure the accuracy of data entry; through the definition of the assigner, the multi-level and multi-dependent field dependency relationships can be automatically calculated according to the value of a certain input field, thereby ensuring that even if there are complex dependencies between multiple data fields, they can be efficiently entered and run correctly. This can significantly promote the electronic management of the laboratory's scientific research data, including data and orders for outsourced experiments, and realize efficient retrieval of scientific research plans, scientific research data and other related content.

[0135] For example, combined with Fig.12, the temperature data and humidity data obtained by the large language model (LLM) are 45 degrees Celsius and 70% respectively, while when the temperature is 45 degrees Celsius, the humidity cannot be 70%. At this time, a prompt can be issued in the display interface, for example, the humidity data can be highlighted to attract the attention of the operator; a sound prompt can also be issued using a buzzer, etc.; or it can be displayed in the form of text below the humidity field to prompt the operator. The embodiment of the present application can be set to verify the data immediately when each data is obtained, and after all the data are obtained, verification is performed again based on the combined verification relationship between the related data to ensure the accuracy of the experiment and the experimental records.

[0136] The embodiment of the present application converts the experimental records of different modalities into data forms (e.g., text) that can be recognized by the large language model (LLM) through the encoder of the corresponding modality. Subsequently, the powerful dialogue and task understanding capabilities of the current large language model (LLM) are used to update and record the experimental records of different modalities. In addition, in the process of dialogue with the large model, scientific research data with other modal recognition errors can also be corrected through interactive inputs such as text and voice. While completing the identification and recording of scientific research data, high-quality data for training and fine-tuning the multimodal scientific research large model is collected.

[0137] In one embodiment of the present application, Fig.16 As shown, the method further includes step S1601 and step S1602:

[0138] In step S1601, based on the storage instruction input by the operator, after the structured scientific research record is stored in the database according to the scientific research unit scheme, the data corresponding to the data field stored in the data field input box of the scientific research unit scheme is released; or,

[0139] In step S1602, based on the non-storage instruction input by the operator, the data corresponding to the data field stored in the data field input box of the scientific research unit program is directly released.

[0140] In an embodiment of the present application, the data in the scientific research unit program data field input box is not permanently stored data, but temporarily stores various data in the scientific research / experimental process. After the final structured scientific research record is generated, the data in the scientific research unit program data field input box will be cleared or migrated to other storage devices. For example, a multimodal scientific research data entry system is installed on a local terminal, and the data stored in the scientific research unit program data field input box is correspondingly stored in the cache of the local terminal. When it is necessary to store it as the final structured scientific research record, the operator clicks the "Submit" button on the display interface, and the current structured scientific research record will be stored in the local database. For another example, a multimodal scientific research data entry system is installed on a cloud server, and the data stored in the scientific research unit program data field input box is correspondingly stored in the cloud server. When it is necessary to store it as the final structured scientific research record, the operator clicks the "Submit" button on the display interface, and the current structured scientific research record will be stored in the cloud database.

[0141] It is worth noting that when the structured scientific research records are stored in the database, based on the data field format specification, a structured storage scheme will be automatically generated for the scientific research unit scheme, so that the operator can submit the structured scientific research records based on the scientific research unit scheme, and each data field should comply with the constraints and data verification relationship of the model. In an embodiment of the present application, whether the data field defined by the operator is a simple text or a more complex multimodal data, the large language model can automatically generate a corresponding data structure for it, and ensure the consistency and integrity of the data during the input and storage process. Since the automatically generated data structure is based on standardized protocols and definitions, scientific research data can be easily shared globally. This automation of structured storage not only simplifies the work of scientific researchers, but also improves the reproducibility of scientific research data and the ability to collaborate across laboratories. In addition, through the automatically generated data structure, scientific researchers can focus on experiments and data records without worrying about underlying data management issues. This feature greatly reduces the data management burden of scientific researchers and improves the efficiency of scientific research work.

[0142] In one embodiment of the present application, Fig.17 As shown, the method further includes step S1701:

[0143] In step S1701, based on the management instruction input by the operator, the content stored in the input box of the scientific research unit program data field is managed.

[0144] The management operation includes at least one of the following operations: deleting data, storing data and modifying data.

[0145] For example, the operator manually deletes, modifies, and performs other operations on the experimental data in the input box of the scientific research unit program data field.

[0146] For another example, when an operator finds a problem with the temperature data in a set of scientific research data, he says "correct the third set of temperature to 45°C". The AI ​​system receives the voice data and can use the built-in algorithm to directly process the voice data and convert the voice data into structured operation instructions:

[0147]

[0148] Then, based on the above structured operation instructions, the AI ​​system will fill in the structured scientific research data (the corresponding scientific research unit plan data field is "temperature" and the value is "45℃") into the corresponding scientific research unit plan data field input box; or, the large language model converts the voice data into unstructured text data through a voice encoder, and then converts the unstructured text data into structured text data, that is, the scientific research unit plan data field is "temperature" and the value is "45℃". Then, the structured data is filled in the corresponding scientific research unit plan data field input box.

[0149] For another example, when the temperature of an experiment is entered incorrectly and needs to be modified, the operator says "change the third set of temperature to 45°C". After the AI ​​system changes the value in the scientific research unit program data field input box of the temperature field to 45°C, the AI ​​system can also provide feedback, such as providing natural language text "The third set of temperature has been corrected to 45°C", or further broadcasting the text through voice broadcast, so that the operator can obtain feedback from the AI ​​system through voice without directly viewing the first display area.

[0150] In one embodiment of the present application, Fig.18 As shown, the method further includes step S1801 and step S1802:

[0151] In step S1801, an instruction to record the experimental record in a predetermined national language is obtained;

[0152] In step S1802, the structured operation instruction is translated into the predetermined national language and recorded in the experimental record.

[0153] For example, when the operator's native language is Chinese and the desired predetermined national language is English, the user can interact with the AI ​​system in the language he is good at. For example, he can issue an operation instruction in Chinese: "Please record the weather as sunny" (the scientific research unit program contains a data field with an ID of "weather"). The AI ​​system can automatically understand the user's operation instruction and generate the following structured operation instruction:

[0154] {

[0155] "operation":"update",

[0156] "field_id":"weather",

[0157] "field_value":"Sunny"

[0158] }

[0159] Since the national language set by the operator is English, the AI ​​system can further translate the information in the above operation instructions into English ("sunny" is translated into "Sunny"), and fill in "Sunny" in the data field input box corresponding to "weather".

[0160] The embodiment of the present application can support the input of multimodal scientific research data in the question-and-answer process, and conduct the question-and-answer process based on non-text modal data such as pictures, videos, or voices, thereby expanding the use scenarios of the question-and-answer process. In addition, by expressing information in different forms through different modal data, the expression of information is richer, thereby being able to provide richer information to the large model, helping to improve the accuracy of the large model's understanding of the input content, thereby generating more accurate reply content and improving the accuracy of the reply content. In the question-and-answer process, when the conversation content includes non-text content, the method of the embodiment of the present application performs multimodal intent recognition on the conversation content and the non-text content. The multimodal intent recognition result indicates the content relevance between the conversation content and the non-text content, so that the corresponding target reply model can be selected based on the content relevance, thereby being able to improve the accuracy of the selected reply model, and then improve the accuracy of the reply content.

[0161] Combine the following Fig.19 The present invention is again described in detail by taking the AI ​​model as an example.

[0162] like Fig.19 As shown, it is a schematic diagram of the interaction between the operator, the AI ​​interaction front end, the data recording end, the AI ​​back end and the AI ​​database.

[0163] Depend on Fig.19 It can be seen that the data in the embodiment of the present application is mainly interacted between the operator, the AI ​​interaction front end, the data recording end, the AI ​​back end and the AI ​​database. The input method can be divided into three main stages: the first stage is the data processing stage, the second stage is the output stage, and the third stage is the stage of saving the interaction information of the operator / AI.

[0164] In the first stage, the input multimodal scientific research data is sent to the AI ​​backend for processing to generate structured operation instructions, that is, multiple operation instructions, and execute multiple operation instructions to obtain data records.

[0165] Here are the steps:

[0166] Step 1: Input the relevant information of the data field (DFs) to be recorded in the AI ​​interaction front end in the form of dialogue / voice / picture, etc.

[0167] Step 2: The AI ​​interactive front end sends the operator's instructions to the AI ​​back end.

[0168] Step 3: Based on the scientific research unit scheme and data field format specification (data field JSON Schema), AI processes the operator's front-end input information into structured operation instructions in the form of a list on the data record side.

[0169] Step 4: The AI ​​backend sends the operation instructions to the data recording end.

[0170] Step 5a: The data recording end processes each operation instruction one by one and generates Acknowledge information corresponding to each operation instruction.

[0171] Step 5b: Every time the data recording end successfully processes an operation instruction, the data field input box filled in by the operation instruction and the filled value are reflected in the data recording interface in real time.

[0172] In the second stage, AI can also generate AI second output. The steps are as follows:

[0173] Step 6a: The data recording end transmits the tabular Acknowledge information to the AI ​​backend.

[0174] Step 6b, the AI ​​backend generates the AI ​​second output, i.e., the reply to the operator, based on a. the original input of the operator, b. the AI ​​first output (operation instruction list), and c. the response of the data recording end (Acknowledge information list).

[0175] Step 6c, the AI ​​backend sends the AI ​​second output to the AI ​​interaction frontend.

[0176] Step 6d: The AI ​​interactive front end displays the information of the previous step to the operator (usually in the form of a dialogue answer)

[0177] In the third stage, the data records can be saved to the database, and the content of the multi-round human-computer question and answer can also be saved to the database, where the database can be a local database or a cloud database. The steps are as follows:

[0178] Step 7, the AI ​​backend stores information such as a. the operator's original input, b. the AI ​​first output (operation instruction list), c. the data recording end response (Acknowledge information list), and d. the AI ​​second output into the AI ​​database.

[0179] In an embodiment of the present application, first obtain a handwritten text photo and a device photo (in other embodiments, it can be a video), which can be manually input into the execution subject (for example, in a computer) by an operator. Based on the scientific research unit plan prompt words corresponding to the specific scientific research unit plan, the picture encoder converts the handwritten text photo and the device photo into corresponding descriptive text. A dialogue robot based on the AI ​​system (the dialogue robot can be part of the AI ​​system) outputs scientific research data. These scientific research data are stored in a temporary scientific research unit plan data field input box to facilitate temporary recording and modification of experimental results. After the operator confirms the results, the experimental record results of this experiment will be stored in the data end of the system.

[0180] The use of the scientific research unit program data field input box can flexibly control the output format of the experimental record, and enable the model to modify and manually correct the recorded erroneous results. These manual correction operations will be collected as manual feedback data to provide an instruction data basis for the subsequent training and fine-tuning of the unified joint representation multimodal model.

[0181] It is worth noting that from inputting the handwritten text photo and the device photo into the picture encoder (belonging to the large language model), to the conversation robot based on the large language model (also belonging to the large language model (LLM)) storing the data in the recording slot, these processes can all be completed by the large language model (the various functions mentioned in this application can also be achieved by using the multimodal large model). This can help operators record scientific research data more quickly and reduce the workload of manual input.

[0182] The multimodal scientific research data recording method of the embodiment of the present application can be applied to technical fields such as medical records, legal document management, and education data analysis that require AI system recording and intelligent management. As an example, it can be directly applied to electronic medical record scenarios. Similar to the scientific research described in other embodiments of the present application, in the medical / hospital / clinical scenario, each institution / department has a demand for medical data recording (such as electronic medical records), and the content and type specifications of the data required to be recorded in different departments, different diseases, and different medical scenarios are not the same. It can be imagined that if each records in a customized manner, it will not only be inefficient, but also not conducive to sharing and experience accumulation. With the help of the scientific research activity management and application platform in the embodiment of the present application, operators in different professional fields can customize the protocols, models, and valuers of medical records according to the scientific research unit syntax, tailor-made medical record methods for specific needs, and the customized scientific research unit scheme can be widely shared in different departments, so that one-time design can be achieved, and the whole hospital / cross-hospital application can be achieved to achieve the unification and standardization of specific medical records. In addition, when recording clinical medical data, it often occurs in the doctor-patient consultation scenario that the doctor may not be able to free his hands to operate a computer or other equipment / handwrite to record clinical medical data during the consultation process. The multimodal data (e.g., voice-based) recording solution proposed in this application provides a convenient solution to the recording needs of this scenario. In summary, this application has huge application potential and significant economic benefits in multiple industries.

[0183] The multimodal scientific research data recording method of the embodiment of the present application converts the multimodal scientific research data of the current scientific research unit plan into structured operation instructions, and then combines at least one round of question and answer process with the operator to convert the structured operation instructions into structured scientific research records according to the scientific research unit plan. This can reduce manual input time and improve experimental recording efficiency. In addition, the AI ​​system is used to automatically identify and record multimodal scientific research data, and errors are corrected through dialogue interaction, thereby ensuring data accuracy and consistency.

[0184] like Fig. 20 As shown, it is a schematic block diagram of a multimodal scientific research data entry system 2000 according to an embodiment of the present application. Fig. 20 As shown, the multimodal scientific research data entry system 2000 of the embodiment of the present application includes an interactive front-end module 2001 and a back-end module 2002. Among them,

[0185] The backend module 2002 is used to receive multimodal scientific research data collected by at least one input device through an input interface based on the current scientific research unit plan.

[0186] Among them, the multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data and structured data.

[0187] The back-end module 2002 is also used to perform structured processing on the multimodal scientific research data based on data field format specifications, so as to convert the multimodal scientific research data into structured operation instructions; and convert the data into structured scientific research records according to the scientific research unit plan.

[0188] The structured operation instruction may be composed of multiple operation instructions, and structured data may be generated when the multiple operation instructions are executed.

[0189] The interactive front-end module 2001 includes a first display area and a second display area, the first display area is used to display and / or play unstructured data, and the second display area is used to display structured data.

[0190] The embodiment of the present application does not require manual filling, but all experimental records are automatically filled in by the AI ​​system, which improves efficiency compared to the traditional manual filling method.

[0191] Combine the following Fig.21 The terminal supporting multimodal scientific research data entry of the present application is described, wherein: Fig.21 A schematic block diagram of a terminal supporting multimodal scientific research data entry according to an embodiment of the present application is shown.

[0192] like Fig.21 As shown, the terminal 2100 supporting multimodal scientific research data entry includes an input device 2101. The input device 2101 may include at least one of the following: a keyboard, a mouse, a microphone, a scanner, a camera, a handwriting board, a drawing board, a USB interface, a network card, and other devices.

[0193] Continue to combine Fig.21 The terminal 2100 supporting multimodal scientific research data entry may also include: one or more memories 2102 and one or more processors 2103, wherein the memory 2102 stores a computer program executed by the processor 2103, and when the computer program is executed by the processor 2103, the processor 2103 executes the multimodal scientific research data recording method described above.

[0194] The terminal 2100 supporting multimodal scientific research data entry may be part or all of a computer device that can implement a multimodal scientific research data recording method through software, hardware, or a combination of software and hardware.

[0195] like Fig.21As shown, the terminal 2100 supporting multimodal scientific research data entry includes one or more memories 2102, one or more processors 2103, a display (not shown) and a communication interface, etc. These components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). It should be noted that Fig.21 The components and structure of the terminal 2100 supporting multimodal scientific research data entry are merely exemplary and non-restrictive. The terminal 2100 supporting multimodal scientific research data entry may also have other components and structures as needed.

[0196] The memory 2102 is used to store various data and executable program instructions generated during the operation of the related method, such as for storing various application programs or algorithms for implementing various specific functions. It can include one or more computer program products, and the computer program product can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc.

[0197] The processor 2103 can be a central processing unit (CPU), a graphics processing unit (GPU), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can be other components in the terminal 2100 that supports multimodal scientific research data entry to perform desired functions.

[0198] In one example, the terminal 2100 supporting multimodal scientific research data entry also includes an output device that can output various information (such as images or sounds) to the outside (such as an operator), and can include one or more of a display device, a speaker, etc.

[0199] The communication interface can be an interface of any currently known communication protocol, such as a wired interface or a wireless interface, wherein the communication interface may include one or more serial ports, USB interfaces, Ethernet ports, WiFi, wired networks, DVI interfaces, device integrated interconnect modules or other suitable ports, interfaces, or connections.

[0200] In addition, according to an embodiment of the present application, a storage medium is also provided, on which program instructions are stored, and when the program instructions are run by a computer or a processor, the corresponding steps of the multimodal scientific research data recording method of the embodiment of the present application are executed. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.

[0201] In addition, an embodiment of the present application also provides a computer program product, which implements the steps of the above method when the computer program / instructions are executed by a processor.

[0202] The system supporting multimodal scientific research data entry, the terminal supporting multimodal scientific research data entry, the storage medium and the computer program product of the embodiments of the present application have the same advantages as the aforementioned multimodal scientific research data recording method because they can implement the aforementioned multimodal scientific research data recording method.

[0203] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present application to this. Those of ordinary skill in the art may make various changes and modifications therein without departing from the scope and spirit of the present application. All these changes and modifications are intended to be included within the scope of the present application as required by the appended claims.

[0204] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0205] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.

[0206] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.

[0207] Similarly, it should be understood that in order to streamline the present application and help understand one or more of the various inventive aspects, in the description of the exemplary embodiments of the present application, the various features of the present application are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present application should not be interpreted as reflecting the following intention: the claimed application requires more features than the features clearly stated in each claim. More specifically, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with features less than all the features of a single disclosed embodiment. Therefore, the claims following the specific embodiment are hereby explicitly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present application.

[0208] It will be understood by those skilled in the art that, except for mutually exclusive features, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this specification may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature that provides the same, equivalent or similar purpose.

[0209] In addition, those skilled in the art will appreciate that, although some embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present application and form different embodiments. For example, in the claims, any one of the claimed embodiments can be used in any combination.

[0210] The various component embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all functions of some modules according to the embodiments of the present application. The application can also be implemented as a device program (e.g., computer program and computer program product) for executing a part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0211] It should be noted that the above embodiments illustrate the present application rather than limit the present application, and that those skilled in the art may design alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets should not be constructed as a limitation to the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of multiple such elements. The present application may be implemented by means of hardware including several different elements and by means of a suitably programmed computer. In a unit claim that lists several devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc. does not indicate any order. These words may be interpreted as names.

[0212] The above is only a specific implementation or description of a specific implementation of the present application, and the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. The protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A multimodal scientific research data recording method, characterized in that: Applied to a multimodal scientific research data entry system, the multimodal scientific research data entry system includes a backend module; the method includes: Based on the current scientific research unit scheme, receiving multimodal scientific research data collected by an input device through an input interface; wherein the multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data and structured data; Based on the data field format specification, calling the back-end module to perform structured processing on the multimodal scientific research data, so as to convert the multimodal scientific research data into structured operation instructions; The structured operation instructions are converted into structured scientific research records according to the scientific research unit plan.

2. The method according to claim 1, characterized in that: Based on the data field format specification, calling the back-end module to perform structured processing on the multimodal scientific research data so as to convert the multimodal scientific research data into structured operation instructions, including: Identifying and classifying the multimodal scientific research data by a data modality identification module to determine the type of the multimodal scientific research data; In the case where the multimodal scientific research data is unstructured data, based on the data type of the multimodal scientific research data, an encoder corresponding to the data type of the multimodal scientific research data in the back-end module is called to convert the multimodal scientific research data into the structured operation instructions.

3. The method according to claim 2, characterized in that The multimodal scientific research data is voice data; the method further comprises: Invoking a speech encoder corresponding to the speech data so that the speech encoder converts the speech data into unstructured text data; The unstructured text data is output.

4. The method according to claim 2, characterized in that: The multimodal scientific research data is image data; the method further comprises: Calling an image encoder corresponding to the image data, so that the image encoder converts the image data into unstructured text data through a text recognition algorithm and / or an optical character recognition algorithm based on the annotated industry knowledge data set; The unstructured text data is output.

5. The method according to claim 1, characterized in that in, The text data is unstructured text data.

6. The method according to any one of claims 3 to 5, characterized in that: The multimodal scientific research data entry system further includes an interactive front-end module; the interactive front-end module includes a first display area and a second display area; the method further includes: Sending the unstructured text data to the first display area so that the first display area displays the unstructured text data in the form of a text block; and / or Receive an unstructured text instruction input by an operator in the first display area, and respond to the unstructured text instruction.

7. The method according to claim 6, characterized in that The multimodal scientific research data is structured data; the method comprises: Sending the structured text data to the second display area so that the second display area displays the structured text data; Among them, the data type and value of the structured data should follow the customized rules of the scientific research unit program data field input box.

8. The method according to claim 6, characterized in that The method further comprises: Playing the unstructured text data in a voice interactive manner; and / or Receive the voice command of the operator and respond to the voice command.

9. The method according to claim 6, characterized in that Convert the unstructured text instructions into structured operation instructions according to the scientific research unit plan, including: Based on the current scientific research unit program prompt words, a preset algorithm is used to analyze and understand the unstructured text instructions to convert the unstructured text data into corresponding structured operation instructions.

10. The method according to claim 9, characterized in that Based on the current scientific research unit program prompt word, the unstructured text instruction is analyzed and understood by using a preset algorithm to convert the unstructured text instruction into a corresponding structured operation instruction, further comprising: An embedded tool specific to the scientific research unit program prompt word architecture is used to vectorize each knowledge point in each text block, and store it in the form of key-value pairs in the scientific research unit program data field input box for subsequent fast matching indexing.

11. The method according to claim 10, characterized in that in, The scientific research unit plan prompt words are preset in the back-end module in advance; and / or, the scientific research unit plan prompt words are automatically created by the back-end module according to the knowledge base corresponding to the current scientific research unit plan.

12. The method according to claim 7, characterized in that Convert the structured operation instructions into structured scientific research records according to the scientific research unit plan, including: The structured data is stored in the scientific research unit scheme data field input box and displayed in the second display area.

13. The method according to claim 12, characterized in that The method further includes: based on the data field format specification, calling the back-end module to perform structured processing on the multimodal scientific research data so as to convert the multimodal scientific research data into structured operation instructions, including: Calling the backend module and using the built-in algorithm of the backend module to directly convert the multimodal scientific research data into structured data; The structured data is stored in the scientific research unit program data field input box and displayed in the second display area.

14. The method according to claim 12 or 13, characterized in that The method further comprises: Converting the structured operation instruction into a code marked with the structured operation instruction through the back-end module; Execute the code marked with the structured operation instruction to form a structured scientific research experiment record.

15. The method according to claim 14, characterized in that The scientific research unit program data field input box is also used to store multiple data records; the method also includes: Based on the constraint relationship and / or data verification relationship between the multiple data records, the multiple data records are verified, and if there are logical errors in the multiple data records, a prompt indicating the existence of the logical errors is sent.

16. The method according to claim 14, characterized in that The method further comprises: Based on the storage instruction input by the operator, after the experimental record is stored in the data according to the scientific research unit plan, the text record and the text summary stored in the data field input box of the scientific research unit plan are released: or, Based on the non-storage instruction input by the operator, the text record and the text summary stored in the scientific research unit program data field input box are directly released.

17. The method according to claim 16, characterized in that The method further comprises: Based on the management instructions input by the operator, the content stored in the data field input box of the scientific research unit program is managed; The management operation includes at least one of the following operations: deleting data, storing data, and modifying data.

18. The method according to claim 6, characterized in that The first display area and the second display area are respectively provided with scroll bars. When the first display area and / or the second display area displays a lot of content, the scroll bars automatically scroll to display the latest content.

19. The method according to claim 1, characterized in that The method further comprises: Obtaining instructions for recording the experimental records in a predetermined country language; The structured operation instructions are translated into the predetermined national language and recorded in the experimental record.

20. The method according to claim 1, characterized in that The input device includes at least one of the following: a keyboard, a mouse, a microphone, a scanner and a camera.

21. A multimodal scientific research data entry system, characterized in that: The system includes a back-end module and an interactive front-end module; wherein, The back-end module is used to receive multimodal scientific research data collected by at least one input device through an input interface based on the current scientific research unit plan; based on the data field format specification, structure the multimodal scientific research data to convert the multimodal scientific research data into structured operation instructions; convert the structured operation instructions into structured scientific research records according to the scientific research unit plan; wherein the multimodal scientific research data includes at least one of the following data forms: text data, voice data, image data and structured data; The interactive front-end module includes a first display area and a second display area, the first display area is used to display and / or play unstructured data, and the second display area is used to display structured data.

22. A terminal supporting multimodal scientific research data entry, characterized in that: The terminal comprises: At least one input device, wherein the input device collects multimodal scientific research data and sends the multimodal scientific research data to the processor for processing; the input device includes at least one of the following: a keyboard, a mouse, a microphone, a scanner, and a camera; A memory and the processor, wherein the memory stores a computer program executed by the processor, and when the computer program is executed by the processor, the processor executes the multimodal scientific research data recording method as described in any one of claims 1 to 20.

23. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, enables the processor to execute the multimodal scientific research data recording method as described in any one of claims 1 to 20.

24. A computer program product, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 20 are implemented.

Citation Information

Patent Citations

  • Multimode physiologic and behavioral data integration and collection system

    CN107137096A

  • Task solution-oriented training method and use method of generative large language model

    CN116756564A

  • Intelligent adaptive retrieval enhancement system and method and storage medium

    CN118210983A

Cited By

  • Multimodal research data recording method, and system, terminal and storage medium

    WO2026152604A1