Rock mine identification method and related device
By acquiring professional information on well logging rocks and minerals and geological logging, and using low-rank matrix factorization and federated learning techniques to train a multimodal large model, the problem of relying on human experience for well logging rock and mineral image recognition was solved, achieving efficient and accurate rock and mineral identification and information description.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA PETROCHEMICAL CORP
- Filing Date
- 2024-10-17
- Publication Date
- 2026-04-17
AI Technical Summary
Current technologies for identifying rock and mineral logging images rely on the experience of geologists, resulting in low efficiency and unreliable accuracy, which cannot meet the demand for accurate and efficient identification of large-scale rock and mineral logging images.
By acquiring logging rock and mineral information and geological logging professional information, a multimodal large model is trained using low-rank matrix factorization, and then updated using federated learning technology to generate a target multimodal large model, thereby achieving accurate identification and information description of logging rock and mineral images.
It improves the accuracy and efficiency of well logging rock and mineral image recognition, provides professional and accurate information description, reduces the complexity of training data, avoids overfitting problems, and enhances the model's generalization ability and training speed.
Smart Images

Figure CN121884048A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of big data technology, and in particular to a method and apparatus for identifying rocks and minerals. Background Technology
[0002] In the field of geological exploration, accurate identification of logging rock and mineral deposits is necessary to reveal underground structures and the distribution of mineral resources. However, the identification of acquired logging rock and mineral images often relies on the experience and domain knowledge of geologists, resulting in low efficiency and unreliable accuracy.
[0003] Therefore, in order to accurately and efficiently identify rock and mineral images from well logging and provide corresponding information descriptions, a rock and mineral identification method is urgently needed. Summary of the Invention
[0004] In view of this, embodiments of this application provide a rock and mineral identification method and related apparatus, which aim to achieve accurate and efficient identification of logging rock and mineral images and provide corresponding information descriptions.
[0005] In a first aspect, embodiments of this application provide a rock and mineral identification method, the method comprising:
[0006] Obtain logging information on rock and mineral deposits, as well as professional information on geological logging.
[0007] The logging rock and mineral information and the geological logging professional information are processed, and the processed logging rock and mineral information and geological logging professional information are used as training data in the training database;
[0008] On each client, a multimodal large model is trained using low-rank matrix factorization based on the training data; the training gradient information of each client is obtained, and the multimodal large model is integrated and trained based on the training gradient information to obtain the target multimodal large model;
[0009] The target multimodal large model is used to identify the logging rock and mineral images to be identified and generate information descriptions based on the question information.
[0010] Optionally, the logging rock and mineral information includes images of the logging rock and mineral and corresponding descriptive information, the descriptive information including tabular descriptive information, and the method further includes:
[0011] Convert the tabular description information into text description information;
[0012] The process of processing the logging rock and mineral information and the geological logging professional information, and using the processed logging rock and mineral information and geological logging professional information as training data in the training database, includes:
[0013] The textual description information is used as the corresponding tag information for the image;
[0014] The image, the tag information, and the geological logging information are converted into their respective formats, and the converted image, the tag information, and the geological logging information are used as training data in the training database.
[0015] Optionally, the multimodal large model is a language large model, and the training data further includes question information. The step of training the multimodal large model based on the training data using low-rank matrix factorization includes:
[0016] Based on the image after the format conversion, the tag information after the format conversion, the geological logging information after the format conversion, and the question information, the language large model is trained using low-rank matrix factorization.
[0017] Optionally, the step of using the target multimodal large model to identify the logging rock and mineral image to be identified and generate an information description for the question information based on the question information includes:
[0018] The logging image of the rock and mineral deposit to be identified and the query information are input into the target multimodal large model;
[0019] The target multimodal large model is used to identify the logging rock and mineral image to be identified based on the question information;
[0020] Based on the recognition results, an information description is generated for the question information.
[0021] Optionally, the logging rock and mineral information includes images of the logging rock and mineral, and the acquisition of the logging rock and mineral information includes:
[0022] Images of the well-logged rock and mineral are acquired from different orientations and brightness levels to obtain information about the well-logged rock and mineral.
[0023] Optionally, the processing of the logging rock and mineral information and the geological logging professional information includes:
[0024] The target region in the image of the logging rock mineral is masked using the Canny operator.
[0025] Optionally, the method further includes:
[0026] Federated learning techniques are used to update the target multimodal large model.
[0027] Secondly, embodiments of this application provide a rock and mineral identification device, the device comprising: a first acquisition module, a processing module, a training module, a second acquisition module, and an identification module;
[0028] The first acquisition module is used to acquire logging rock and mineral information and geological logging professional information;
[0029] The processing module is used to process the logging rock and mineral information and the geological logging professional information, and use the processed logging rock and mineral information and geological logging professional information as training data in the training database;
[0030] The training module is used to train a multimodal large model on each client using low-rank matrix factorization based on the training data.
[0031] The second acquisition module is used to acquire the training gradient information of each client during the training, and to integrate and train the multimodal large model based on the training gradient information to obtain the target multimodal large model;
[0032] The recognition module is used to identify the logging rock and mineral images to be identified and generate information descriptions for the question information based on the target multimodal large model.
[0033] Thirdly, this application provides an electronic device, the device comprising: a processor, a memory, and a system bus;
[0034] The processor and the memory are connected via the system bus;
[0035] The memory is used to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform the method described in the first aspect.
[0036] Fourthly, embodiments of this application provide a computer storage medium storing code, wherein when the code is executed, a device executing the code implements the method described in any of the first aspects above.
[0037] This application provides a method and related apparatus for rock and mineral identification. When executing the method, firstly, well logging rock and mineral information and geological logging professional information are acquired. Then, the well logging rock and mineral information and geological logging professional information are processed, and the processed well logging rock and mineral information and geological logging professional information are used as training data in a training database. On each client, a multimodal large model is trained using low-rank matrix factorization based on the training data. The training gradient information of each client is obtained, and the multimodal large model is integrated and trained based on the training gradient information to obtain a target multimodal large model. Finally, the target multimodal large model is used to identify the well logging rock and mineral image to be identified and generate an information description corresponding to the question information based on the question information. Thus, by using geological logging professional information as training data for training the multimodal large model, the multimodal large model can understand the context of the well logging rock and mineral data through training, improving the accuracy of identifying geological structures and features, and facilitating the provision of accurate information descriptions for well logging rock and mineral data. Simultaneously, multiple clients are introduced during training using low-rank matrix factorization. The multimodal large model is integrated and trained based on the training gradient information from each client. This allows for collaborative training across different devices and data sources, aggregating knowledge and patterns from different domains and improving the model's generalization ability. It also increases the training speed of the multimodal large model. Furthermore, by reducing the complexity of the training data, less data can be used to display the situation of logging rock and mineral data, avoiding overfitting. Finally, the target multimodal large model, based on the trained data, can identify the logging rock and mineral images to be identified and generate information descriptions corresponding to those questions. This ensures accurate and efficient identification of logging rock and mineral images while guaranteeing the professionalism and accuracy of the provided information descriptions. Attached Figure Description
[0038] To more clearly illustrate the technical solutions in this embodiment or the prior art, the drawings used in the description of the embodiment or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0039] Figure 1 A flowchart illustrating a rock and mineral identification method provided in this application embodiment;
[0040] Figure 2 This application provides schematic diagrams illustrating different forms of logging rock deposits in its embodiments.
[0041] Figure 3 A schematic diagram of an image after masking, provided in an embodiment of this application;
[0042] Figure 4 A schematic diagram illustrating a tabular description of information provided in an embodiment of this application;
[0043] Figure 5 A schematic diagram illustrating rock and mineral identification based on a target polymorphic large model, provided as an embodiment of this application;
[0044] Figure 6 This is a schematic diagram of the structure of a rock and mineral identification device provided in an embodiment of this application. Detailed Implementation
[0045] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0046] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0047] Research on related technologies has revealed that accurate identification of rocks and minerals is crucial in geological exploration for revealing underground structures and the distribution of mineral resources. However, when identifying rock and mineral images acquired from well logging, these technologies often rely on the experience and domain knowledge of geologists, incurring significant labor costs, and the accuracy of image identification cannot be guaranteed. This approach cannot meet the needs of identifying large-scale well logging rock and mineral images, nor can it provide users with timely descriptive information about the images.
[0048] Based on this, this application proposes a rock and mineral identification method and related apparatus. Geological logging information is used as training data to train a multimodal large-scale model, enabling the model to understand the context of the logging rock and mineral data, improving the accuracy of geological structure and feature identification, and facilitating the provision of accurate information descriptions for the logging rock and mineral data. Simultaneously, by utilizing low-rank matrix factorization to train the multimodal large-scale model, the training speed can be improved. Based on the trained target multimodal large-scale model, accurate and efficient identification of logging rock and mineral images is achieved, while ensuring the professionalism and accuracy of the provided information descriptions.
[0049] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0050] Figure 1 A flowchart of a rock and mineral identification method provided in an embodiment of this application is shown below. Figure 1 As shown in the embodiments of this application, a rock and mineral identification method includes:
[0051] S11: Obtain logging information on rock and mineral deposits, as well as geological logging information.
[0052] The logging information may include images of the logging rock and mineral deposits, along with corresponding descriptive information. The acquisition of these images is affected by factors such as lighting, instrument angle, and shooting orientation, resulting in differences between images of the same logging rock and mineral deposit. To increase the richness of training samples and ensure sufficient training of the multimodal large-scale model, during step S11 ("acquiring logging information"), images of the logging rock and mineral deposits are acquired under different orientations and brightness levels. Alternatively, the acquired images are subjected to varying degrees of brightness and angle interference to provide noise levels more closely resembling the actual environment of the logging rock and mineral deposits. This allows the multimodal large-scale model to capture different representations of the logging rock and mineral deposits, thereby fully learning their characteristics. Figure 2 Schematic diagrams illustrating different forms of logging rock deposits provided in embodiments of this application, see [link to relevant documentation]. Figure 2 As shown, the images of the same logging rock mineral exhibit different display effects under varying color saturation and light intensity interference.
[0053] In addition to processing the images of well-logged rock and mineral deposits to enrich the training samples for training the multimodal large model, the Canny operator can be used for region description and mask filling to improve the accuracy of mineral composition analysis in these images. This involves masking the target regions in the image. The target regions mentioned here refer to secondary regions in the image, which can be understood as areas in the image with weak correlation to the well-logged rock and mineral deposits, or parts that are not part of the well-logged rock and mineral deposits. The aforementioned Canny operator is a multi-level edge detection algorithm that ensures the detected edges are as close as possible to the actual edges, detects as many edges as possible, and minimizes noise interference with edge detection.
[0054] Figure 3 A schematic diagram of an image after masking, provided in an embodiment of this application, see [link / reference]. Figure 3 As shown, by detecting and identifying secondary regions in the image based on the Canny operator and then occluding these secondary regions through a masking operation, other regions in the image related to the logging rock and mineral deposits can be highlighted. Thus, when training a multimodal large model using this logging rock and mineral deposit information as training data, the model can focus its learning attention on the primary regions in the image related to the logging rock and mineral deposits, enabling it to learn the features of the logging rock and mineral deposits more efficiently.
[0055] S12: Process the logging rock and mineral information and the geological logging professional information, and use the processed logging rock and mineral information and geological logging professional information as training data in the training database.
[0056] As mentioned above, in this embodiment of the application, the logging rock and mineral information includes images of the logging rock and mineral and corresponding descriptive information. This descriptive information can include textual and tabular descriptive information. Textual information can be directly obtained, while tabular descriptive information needs to be converted from tabular to textual form. Figure 4 This is a schematic diagram illustrating a tabular description of information provided in an embodiment of this application. See also... Figure 4 As shown, the image displays descriptive information such as photo number, rock name, and feature description. After processing the tabular descriptive information, the logging rock and mineral information now includes the image of the logging rock and mineral, along with its corresponding textual descriptive information.
[0057] The aforementioned S12 mentions "processing the logging rock and mineral information and the geological logging professional information, and using the processed logging rock and mineral information and the geological logging professional information as training data in the training database". This method can be as follows: first, using the textual description information as the corresponding image label information, then converting the image, the label information and the geological logging professional information into a format, and using the converted image, the label information and the geological logging professional information as training data in the training database.
[0058] In this application, the geological logging professional information includes, but is not limited to, the following: a large amount of geological logging professional knowledge such as explanations of geological and logging technology terms, rock type classification, and rock characteristic descriptions. By using the geological logging professional information as training data in the training database for training the multimodal large model, the multimodal large model can learn the geological logging professional knowledge related to the logging rocks and minerals, which is beneficial for providing professional information descriptions in response to user questions about the logging rocks and minerals to be identified.
[0059] Textual descriptions are used as labels for the corresponding logging rock and mineral images. Then, the images, labels, and geological logging information are converted to the appropriate formats. This converted content can be used as the training database structure.
[0060] In this application, a training sample in the training database can simultaneously include: an image, label information, question information, and geological logging information. The structure and information of a training sample are illustrated below:
[0061] {"img":"fewshot-data / rock.png"},
[0062] "prompt":"This image was obtained using crossed polarized light. What lithology does it belong to?"
[0063] "label": "This section is composed of granite or monocrystalline feldspar, possibly formed from the fragmentation of skeletal grains in conglomerate. Therefore, it is named granitic conglomerate."
[0064] "history":[["What is crossed polarizer?", "This is the common name for two polarizers that are perpendicular to each other"],["What is conglomerate?", "A rock composed mainly of sandstone grains with different size distributions and low compositional maturity"],...]},
[0065] The `img` keyword represents the path information, which is the image of the logging rock and mineral that needs to be identified. `prompt` represents the question information that needs to be input along with the image. `label` represents the information description that we want to obtain after training the multimodal large model. `history` represents the geological logging professional information that the multimodal large model needs to have before the image input.
[0066] It should be noted that, in the embodiments of this application, when training the multimodal large model, geological logging professional information needs to be input as training data into the multimodal large model for training. When the multimodal large model completes training and obtains the target multimodal large model, when using the target multimodal large model to identify the logging rock and mineral images to be identified, it is not necessary to add geological logging professional information again. It is only necessary to input the logging rock and mineral images and the corresponding question information into the target multimodal large model.
[0067] S13: Train the multimodal large model on each client using low-rank matrix factorization based on the training data;
[0068] S14: Obtain the training gradient information of each client during the training, and perform integrated training on the multimodal large model based on the training gradient information to obtain the target multimodal large model.
[0069] In this embodiment, the multimodal large model can be a language large model. This multimodal large model can learn and receive information through dialogue, thus allowing for the description and interpretation of the acquired logging rock and mineral images using a question-and-answer approach. The training data used to train the multimodal large model also includes question information; that is, a single training data set includes: an image that has undergone the aforementioned format conversion, label information for the aforementioned format conversion, geological logging professional information for the aforementioned format conversion, and question information related to the image.
[0070] The training method in step S13, "training the multimodal large model based on the training data using low-rank matrix factorization under each client", can be as follows: training the language large model under each client using low-rank matrix factorization based on the image after format conversion, the label information after format conversion, the geological logging professional information after format conversion, and the question information.
[0071] Based on the converted images, converted label information, converted geological logging information, and question information, low-rank matrix factorization (LMF) can be used to train the large-scale language model to improve the training speed of the multimodal model and reduce the dimensionality of the weight matrix involved in training. Large-scale multimodal models, employing multi-head attention mechanisms for language and image feature recognition and matching, possess good fitting capabilities and require a large number of training samples. However, this also brings some problems, such as slow training convergence due to the large number of model parameters, resulting in a long training time. Therefore, LMF is used to address these issues. LMF decomposes a high-dimensional matrix into a product of multiple low-dimensional matrices while preserving as many of the main features of the original data as possible.
[0072] This application embodiment utilizes low-rank matrix factorization to reduce the dimensionality of training samples before increasing it, thereby reducing the number of elements in the actual weight matrix and consequently reducing the dimensionality of the weight matrix used in training. This approach improves the training speed of large multimodal models while avoiding overfitting. Furthermore, by combining images, query information, and geological logging information as input data for the large multimodal model, and using textual descriptive information as label information, a training database is constructed, enabling the matching of images and descriptive information.
[0073] In this embodiment of the application, the acquired logging rock and mineral information and geological logging professional information can be divided into training data and test data according to a certain ratio. When the training of the multimodal large model is completed based on the training data, the training multimodal large model is tested using the test data. When the test is passed, the target multimodal large model can be obtained.
[0074] It should also be noted that, in the process of training the target multimodal large model in this embodiment, the multimodal large model can be trained separately on different clients using low-rank matrix factorization based on the training data obtained from each client. Then, the training gradient information of each client is obtained, and the multimodal large model is integrated and trained based on this training gradient information to obtain the target multimodal large model. Here, the training gradient information refers to the data used to guide the adjustment of model parameters during the training process of the multimodal large model.
[0075] The method described above, which integrates and trains a multimodal large model on multiple clients to obtain a target multimodal large model, enables the training process to be carried out across different clients for joint training of the multimodal large model. This facilitates the aggregation of knowledge and models from different clients, thereby improving the generalization ability of the target multimodal large model obtained by integrating the training gradient information from each client for training. This, to a certain extent, ensures the effectiveness of the target multimodal large model in subsequent use.
[0076] S15: Using the target multimodal large model, identify the logging rock and mineral images to be identified based on the question information and generate information descriptions for the question information.
[0077] The implementation method of step S15 can be as follows: First, input the logging rock and mineral image to be identified and the query information into the target multimodal large model. Then, use the target multimodal large model to identify the logging rock and mineral image to be identified based on the query information. Finally, generate an information description for the query information based on the identification result.
[0078] The logging image to be identified, along with the query information for that image, is input into the target multimodal model. The target multimodal model then uses the query information to identify the logging image. Based on the identification results and combined with geological logging information, a professional and accurate information description is generated for the query information. Simultaneously, the identification and information description can be saved as information samples for subsequent updates to the target multimodal model.
[0079] It should also be noted that, in the embodiments of this application, federated learning technology can also be used to update the target multimodal large model in order to improve the model’s performance and meet the ever-evolving recognition requirements.
[0080] When the target multimodal large model needs to be updated, the server first needs to be initialized and the target multimodal large model needs to be sent to all clients. Each client uses its local data D. k Train the target multimodal large model and update the model parameters. Local model training on client k can be represented as:
[0081]
[0082] Among them, w (t) These are the target multimodal large model parameters for the current round, and η is the learning rate. It is the gradient of the loss function. F k This represents the large-scale deep learning network model where the k-th client resides.
[0083] The server collects updated model parameters from each client. k (t+1) The target polymorphic large model is then updated using a weighted average. The weighted average formula is:
[0084]
[0085] Where, n k is the number of data samples for client k, and n is the total number of data samples for all clients.
[0086] Figure 5 A schematic diagram illustrating rock and mineral identification based on a target polymorphic large model, provided as an embodiment of this application, is shown below. Figure 5 As shown, by inputting the image of the logging rock and the question "Describe this rock to me, the yellow area does not need to be identified" into the target multimodal large model, the corresponding descriptive information can be obtained: "The conglomerate in the image is composed of granite and feldspar. Granite grains are distributed throughout the rock... The conglomerate has a rough and irregular shape." In this embodiment, gradio (a Python library for creating interactive interfaces for machine learning models) can be used as a graphical tool to display the question-and-answer effect.
[0087] This embodiment proposes a method for rock and mineral identification. The method first acquires well logging rock and mineral information and geological logging professional information. Then, it processes this information, using it as training data in a training database. Based on this training data, a multimodal large model is trained using low-rank matrix factorization to obtain a target multimodal large model. Finally, the target multimodal large model is used to identify the well logging rock and mineral image to be identified and generate an information description corresponding to the question information. Thus, by using geological logging professional information as training data for the multimodal large model, it can understand the context of the well logging rock and mineral data through training, improving the accuracy of identifying geological structures and features, and providing accurate information descriptions for the well logging rock and mineral data. Furthermore, using low-rank matrix factorization to train the multimodal large model improves the training speed, and by reducing the complexity of the training data, less data can be used to display the situation of the well logging rock and mineral data, avoiding overfitting. Finally, the target polymorphic large model obtained after training can identify the logging rock and mineral images to be identified based on the question information, and generate information descriptions for the question information. It can meet the requirements of accurate and efficient identification of logging rock and mineral images, while ensuring the professionalism and accuracy of the information descriptions provided.
[0088] Figure 6This is a schematic diagram of the structure of a rock and mineral identification device provided in an embodiment of this application, as shown below. Figure 6 As shown, a rock and mineral identification device specifically includes: a first acquisition module 100, a processing module 200, a training module 300, a second acquisition module 400, and an identification module 500;
[0089] The first acquisition module 100 is used to acquire logging rock and mineral information and geological logging professional information;
[0090] The processing module 200 is used to process the logging rock and mineral information and the geological logging professional information, and use the processed logging rock and mineral information and geological logging professional information as training data in the training database;
[0091] The training module 300 is used to train a multimodal large model on each client using low-rank matrix factorization based on the training data.
[0092] The second acquisition module 400 is used to acquire the training gradient information of each client during the training, and to perform integrated training on the multimodal large model based on the training gradient information to obtain the target multimodal large model;
[0093] The recognition module 500 is used to recognize the logging rock and mineral image to be recognized and generate an information description for the question information based on the question information using the target multimodal large model.
[0094] In a possible implementation, the device further includes a conversion module. The logging rock and mineral information includes images of the logging rock and mineral and corresponding descriptive information for the images. The descriptive information includes tabular descriptive information. The conversion module is used for:
[0095] Convert the tabular description information into text description information;
[0096] The processing module 200 is configured to: use the textual description information as the corresponding tag information of the image;
[0097] The image, the tag information, and the geological logging information are converted into their respective formats, and the converted image, the tag information, and the geological logging information are used as training data in the training database.
[0098] In a possible implementation, the multimodal large model is a language large model, the training data also includes question information, and the training module 300 is used for:
[0099] Based on the image that has undergone the format conversion, the tag information that has undergone the format conversion, the geological logging information that has undergone the format conversion, and the question information, the language model is trained on each client using low-rank matrix factorization.
[0100] In a possible implementation, the identification module 500 is used for:
[0101] The logging image of the rock and mineral deposit to be identified and the query information are input into the target multimodal large model;
[0102] The target multimodal large model is used to identify the logging rock and mineral image to be identified based on the question information;
[0103] Based on the recognition results, an information description is generated for the question information.
[0104] In a possible implementation, the logging rock and mineral information includes images of the logging rock and mineral, and the first acquisition module 100 is used for:
[0105] Images of the well-logged rock and mineral are acquired from different orientations and brightness levels to obtain information about the well-logged rock and mineral.
[0106] In a possible implementation, the processing module 200 is used to:
[0107] The target region in the image of the logging rock mineral is masked using the Canny operator.
[0108] In a possible implementation, the device further includes an update module, the update module being configured to:
[0109] Federated learning techniques are used to update the target multimodal large model.
[0110] This embodiment proposes a device for rock and mineral identification, comprising: a first acquisition module, a processing module, a training module, a second acquisition module, and an identification module. The first acquisition module acquires logging rock and mineral information and geological logging information. The processing module processes the logging rock and mineral information and the geological logging information, using the processed information as training data in a training database. The training module trains a multimodal large model on each client using low-rank matrix factorization based on the training data. The second acquisition module acquires the training gradient information from each client and integrates and trains the multimodal large model based on the training gradient information to obtain a target multimodal large model. The identification module uses the target multimodal large model to identify the logging rock and mineral image to be identified based on a query and generates an information description corresponding to the query. Thus, by using geological logging information as training data for training a multimodal large-scale model, the model can understand the context of logging rock and mineral data through training, improving the accuracy of identifying geological structures and features, and providing accurate information descriptions for the logging rock and mineral data. Simultaneously, using low-rank matrix factorization to train the multimodal large-scale model can improve training speed, and by reducing the complexity of the training data, less data can be used to display the situation of logging rock and mineral data, avoiding overfitting. Finally, the target multimodal large-scale model, based on the trained data, can identify the logging rock and mineral images to be identified and generate information descriptions corresponding to those questions, achieving accurate and efficient identification of logging rock and mineral images while ensuring the professionalism and accuracy of the provided information descriptions.
[0111] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatuses and methods according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0112] This application also provides corresponding devices and computer-readable storage media for implementing the solutions provided in this application.
[0113] The device includes a memory and a processor. The memory is used to store instructions or code, and the processor is used to execute the instructions or code to enable the device to perform a rock and mineral identification method according to any embodiment of this application.
[0114] In practical applications, the computer-readable storage medium can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0115] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0116] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0117] Computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0118] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0119] The above description is merely one specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for identifying rocks and minerals, characterized in that, The method includes: Obtain logging information on rock and mineral deposits, as well as professional information on geological logging. The logging rock and mineral information and the geological logging professional information are processed, and the processed logging rock and mineral information and geological logging professional information are used as training data in the training database; The multimodal large model is trained on each client using low-rank matrix factorization based on the training data. The training gradient information of each client during the training is obtained, and the multimodal large model is integrated and trained based on the training gradient information to obtain the target multimodal large model; The target multimodal large model is used to identify the logging rock and mineral images to be identified and generate information descriptions based on the question information.
2. The method according to claim 1, characterized in that, The logging rock and mineral information includes images of the logging rock and mineral and corresponding descriptive information for the images. The descriptive information includes tabular descriptive information. The method further includes: Convert the tabular description information into text description information; The process of processing the logging rock and mineral information and the geological logging professional information, and using the processed logging rock and mineral information and geological logging professional information as training data in the training database, includes: The textual description information is used as the corresponding tag information for the image; The image, the tag information, and the geological logging information are converted into their respective formats, and the converted image, the tag information, and the geological logging information are used as training data in the training database.
3. The method according to claim 2, characterized in that, The multimodal large model is a language large model, and the training data also includes question information. The training of the multimodal large model on each client using low-rank matrix factorization based on the training data includes: Based on the image that has undergone the format conversion, the tag information that has undergone the format conversion, the geological logging information that has undergone the format conversion, and the question information, the language model is trained on each client using low-rank matrix factorization.
4. The method according to claim 1, characterized in that, The step of using the target multimodal large model to identify the logging rock and mineral images to be identified and generate information descriptions for the question information includes: The logging image of the rock and mineral deposit to be identified and the query information are input into the target multimodal large model; The target multimodal large model is used to identify the logging rock and mineral image to be identified based on the question information; Based on the recognition results, an information description is generated for the question information.
5. The method according to claim 1, characterized in that, The logging rock and mineral information includes images of the logging rock and mineral, and the acquisition of the logging rock and mineral information includes: Images of the well-logged rock and mineral are acquired from different orientations and brightness levels to obtain information about the well-logged rock and mineral.
6. The method according to claim 5, characterized in that, The processing of the logging rock and mineral information and the geological logging professional information includes: The target region in the image of the logging rock mineral is masked using the Canny operator.
7. The method according to claim 1, characterized in that, The method further includes: Federated learning techniques are used to update the target multimodal large model.
8. A rock and mineral identification device, characterized in that, The device includes: a first acquisition module, a processing module, a training module, a second acquisition module, and a recognition module; The first acquisition module is used to acquire logging rock and mineral information and geological logging professional information; The processing module is used to process the logging rock and mineral information and the geological logging professional information, and use the processed logging rock and mineral information and geological logging professional information as training data in the training database; The training module is used to train a multimodal large model on each client using low-rank matrix factorization based on the training data. The second acquisition module is used to acquire the training gradient information of each client during the training, and to integrate and train the multimodal large model based on the training gradient information to obtain the target multimodal large model; The recognition module is used to identify the logging rock and mineral images to be identified and generate information descriptions for the question information based on the target multimodal large model.
9. An electronic device, characterized in that, The device includes: a processor, a memory, and a system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform the rock and mineral identification method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an implementation program for the rock and mineral identification method, which, when executed by a processor, implements the steps of the method as described in any one of claims 1-7.