Part identification model training method, part identification method and electronic equipment

By using the component recognition model trained by multimodal information, the problem of low accuracy in foreign object recognition in liquids in nuclear power plants is solved, and more accurate foreign object type judgment and evaluation decisions are achieved.

CN120030343APending Publication Date: 2025-05-23LINGAO NUCLEAR POWER +3
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411975055.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

Due to the presence of water in the pipeline in the nuclear power plant, it is difficult for inspectors to accurately identify the type of foreign objects, which in turn affects foreign objects evaluation decisions.

Method used

Multimodal information is used, including image, text information, video information and audio information in the liquid, to construct the annotated data set, and a training data set is generated through the instruction template to train the component recognition model to obtain the target component recognition model.

Benefits of technology

It improves the accuracy of component identification in liquids in nuclear power plants, enhances accurate judgment of foreign object types, and supports more accurate foreign object evaluation decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030343A_ABST
    Figure CN120030343A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of nuclear power plant maintenance, and provides a part identification model training method, a part identification method and electronic equipment, and the method comprises the steps: obtaining the multi-modal information of a part of a nuclear power plant, the multi-modal information comprises an image in liquid and at least one of the following information: text information, video information and audio information; constructing a data set according to the marked multi-modal information, wherein the marked multi-modal information comprises the type of the part; generating at least one instruction template; constructing an instruction data training set according to each instruction template and the data set; and training a to-be-trained part recognition model according to the instruction data training set to obtain a target part recognition model. Through the method, the accuracy of identifying the types of the parts in the liquid of the nuclear power plant by the target part identification model obtained through training can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of nuclear power plant maintenance technology, and in particular relates to a component recognition model training method, a component recognition method, a component recognition model training device, a component recognition device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] A nuclear power plant, also known as a nuclear power plant, is a facility that uses a nuclear reactor to convert nuclear energy into electrical energy. It usually includes a nuclear reactor, a cooling system, a steam generator, a turbine, a generator, a safety system and other units. Pipes are usually used to transport materials within a unit or between units. For example, pipes are used in the cooling system to circulate coolant to keep the temperature of the reactor core within a safe range.

[0003] With the long-term operation of the unit, some parts may accidentally fall into the pipeline. When foreign objects fall into the pipeline, they may affect the safe and stable operation of the nuclear power plant. However, due to the complex piping system of nuclear power plant equipment and the narrow internal space of some parts, in many cases they cannot be removed without affecting the operation of the nuclear power plant. If the unit is directly shut down to remove the foreign objects, it will cause great losses. Therefore, when it is impossible to directly remove the foreign objects in the pipeline, an endoscope is generally used to deeply check the foreign objects inside the pipeline, and then it is judged whether it is necessary to control the shutdown to remove the foreign objects based on the inspection results.

[0004] However, in nuclear power plants, due to the presence of water in the pipes and the limited visibility of the water, it is difficult for inspectors to determine what maintenance parts the foreign object is, which makes foreign object assessment decisions difficult. Summary of the invention

[0005] The embodiments of the present application provide a component recognition model training method, a component recognition method and an electronic device, which can solve the problem of low accuracy in identifying liquid components in nuclear power plants.

[0006] In a first aspect, an embodiment of the present application provides a component recognition model training method, which is applied to a nuclear power plant, comprising:

[0007] Acquire multimodal information of components of a nuclear power plant, the multimodal information comprising an image in liquid and at least one of the following information: text information, video information, and audio information, wherein the image in liquid is an image of the component in liquid;

[0008] Constructing a data set according to the annotated multimodal information, wherein the annotated multimodal information includes the type of the component;

[0009] Generate at least one instruction template, where the instruction template is used to describe a task that needs to be completed by the component recognition model;

[0010] Constructing an instruction data training set according to each of the instruction templates and the data set;

[0011] The component recognition model to be trained is trained according to the instruction data training set to obtain a target component recognition model.

[0012] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:

[0013] In the embodiment of the present application, an instruction data training set is constructed according to each instruction template and data set, and then the component recognition model to be trained is trained according to the instruction data training set to obtain a target component recognition model. Since the data set is constructed according to multimodal information that at least marks the types of components of the nuclear power plant, the instruction data training set also includes multimodal information that marks the types of components of the nuclear power plant, and compared with the single modal information, the multimodal information includes more contextual information of the nuclear power plant, and the multimodal information includes images in the liquid, so that when the component recognition model to be trained is trained according to the instruction data training set, it is beneficial to improve the accuracy of the target component recognition model obtained by training to identify the types of components in the liquid of the nuclear power plant. In addition, since the instruction template is used to describe the tasks that the component recognition model needs to complete, when the component recognition model to be trained is trained according to the instruction data training set, it is equivalent to training the ability of the component recognition model to be trained to recognize tasks, which is beneficial to improve the accuracy of the target component recognition model in identifying according to the task.

[0014] In a second aspect, an embodiment of the present application provides a component identification method, which is applied to a nuclear power plant, comprising:

[0015] Obtain an image to be recognized;

[0016] The target component recognition model as described in the first aspect is used to recognize the image to be recognized to obtain a recognition result.

[0017] In a third aspect, an embodiment of the present application provides a component recognition model training device, which is applied to a nuclear power plant, comprising:

[0018] A multimodal information acquisition module, used to acquire multimodal information of components of a nuclear power plant, wherein the multimodal information includes an image in liquid and at least one of the following information: text information, video information, and audio information, wherein the image in liquid is an image of the component in liquid;

[0019] A data set construction module, used to construct a data set according to the annotated multimodal information, wherein the annotated multimodal information includes the type of the component;

[0020] An instruction template generation module, used to generate at least one instruction template, wherein the instruction template is used to describe the task that the component recognition model needs to complete;

[0021] An instruction data training set construction module, used to construct an instruction data training set according to each of the instruction templates and the data set;

[0022] The target component recognition model generation module is used to train the component recognition model to be trained according to the instruction data training set to obtain the target component recognition model.

[0023] In a fourth aspect, an embodiment of the present application provides a component identification device, which is applied to a nuclear power plant, including:

[0024] An image acquisition module to be identified, used to acquire an image to be identified;

[0025] The recognition result generating module is used to recognize the image to be recognized by using the target component recognition model as described in the second aspect to obtain a recognition result.

[0026] In a fifth aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the method described in the first aspect is implemented, or the method described in the second aspect is implemented.

[0027] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, and when the computer program is executed by a processor, it implements the method as described in the first aspect, or implements the method as described in the second aspect.

[0028] In a seventh aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device executes the method described in the first aspect or the second aspect above.

[0029] It can be understood that the beneficial effects of the second to seventh aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.

[0031] Figure 1It is a flowchart of a component recognition model training method provided in one embodiment of the present application;

[0032] Figure 2 It is a flowchart of a component identification method provided in one embodiment of the present application;

[0033] Figure 3 It is a structural schematic diagram of a component recognition model training device provided by an embodiment of the present application;

[0034] Figure 4 It is a structural schematic diagram of a component identification device provided in one embodiment of the present application;

[0035] Figure 5 It is a structural schematic diagram of an electronic device provided in another embodiment of the present application. DETAILED DESCRIPTION

[0036] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0037] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.

[0038] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0039] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0040] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the phrases "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. appearing in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways.

[0041] A nuclear power plant is a facility that converts nuclear energy into electrical energy through the cooperation of multiple units. Since pipes are usually used to transfer materials within a unit or between units, and there are many parts on the pipes, after the unit has been in operation for a long time, the parts on the pipes may fall into the pipes. At this time, the parts that fall into the pipes are foreign objects relative to the pipes. In addition, during the overhaul of a nuclear power plant (an overhaul of the unit's equipment that requires the reactor to be shut down), due to the large number of maintenance activities, although all nuclear power plants have implemented strict maintenance quality control, it is difficult to prevent all maintenance activities from completely introducing items outside the system (such as miscellaneous parts such as screws and nuts that are dismantled and replaced) into the system, and these items may also fall into the pipes.

[0042] Preventing foreign matter is very important for the safe and stable operation of nuclear power plants. These foreign matter may block pipelines and affect the flow rate of fluids, or get stuck inside mechanical parts and affect the normal operation of equipment. In particular, foreign matter in the reactor core parts poses a hidden danger to the safe operation of the core parts. For example, the wall thickness of the steam generator heat transfer tube is only 1.09 mm. If metal foreign matter exists between the tube bundles and repeatedly rubs the heat transfer tube, it may cause the thickness of the heat transfer tube wall to thin or even suddenly rupture, thereby causing a loss of coolant accident (LOCA) in the primary circuit. For another example, special material foreign matter that exists in the system for a long time may also gradually release chemical substances, such as silver (Ag). These chemical substances affect the turbidity of the primary circuit liquid, and then affect the chemical indicators of the nuclear power plant.

[0043] When foreign matter is found, the first thing a nuclear power plant needs to consider is to remove it. However, due to the complex piping system of nuclear power plant equipment and the narrow internal space of some components, it is often impossible to remove them with tools. If foreign matter cannot be removed, the nuclear power plant needs to conduct a technical assessment, that is, to assess whether the presence of such foreign matter in the system pipeline for one fuel cycle is within the safety margin, or whether the possible impact on surrounding components is within an acceptable range.

[0044] In order to conduct this technical assessment, we first need to find out what kind of component or item the foreign object is. To this end, an endoscope is generally used to look deep inside the pipe to check the foreign object. However, due to the limited underwater visibility, partial obstruction by surrounding components, and the limited accuracy of the endoscope video, it is difficult for inspectors to determine what kind of component the foreign object is with the naked eye, which makes it difficult to make decisions on foreign object assessment.

[0045] In order to improve the accuracy of foreign body recognition in nuclear power plants, the present application provides a method for training a component recognition model. In this method, the component recognition model is trained using multimodal information including labeled images of nuclear power plant components in liquid and instruction templates to obtain a target component recognition model.

[0046] The component recognition model training method provided in the embodiment of the present application is described below with reference to the accompanying drawings.

[0047] Figure 1 A schematic flow chart of a component identification model training method provided in an embodiment of the present application is shown, which is applied to a nuclear power plant and is described in detail as follows:

[0048] S11, obtaining multimodal information of components of a nuclear power plant, the multimodal information comprising images in liquid and at least one of the following information: text information, video information and audio information, wherein the images in liquid are images of the components in liquid.

[0049] Among them, the parts here include one or more of valve gaskets, screws, nuts, etc.

[0050] Wherein, the above-mentioned liquid includes a coolant, and the coolant includes water.

[0051] The text information refers to document information containing text description information of parts; the video information refers to video information containing image description information of parts; and the audio information refers to document information containing audio description information of parts.

[0052] In the embodiment of the present application, the multimodal information of the parts can be obtained from the maintenance report, experience feedback report and other documents or multimedia materials of the nuclear power plant, that is, from the existing materials. Of course, it can also be regenerated when needed. For example, when it is necessary to obtain an image in a liquid, the parts can be placed in the liquid and then photographed. When it is necessary to obtain text information, it can be obtained by describing the parts in text. The process of obtaining video information and audio information is similar and will not be described here.

[0053] In an embodiment of the present application, since the acquired information includes not only the image of the component in the liquid (i.e., the image in the liquid), but also at least one other information, such as text information, etc., and compared with unimodal information, multimodal information contains more information. Therefore, it is beneficial to obtain more contextual information of the component based on the acquired multimodal information.

[0054] S12, constructing a data set according to the annotated multimodal information, where the annotated multimodal information includes the types of components.

[0055] The annotation here includes manual annotation and automatic annotation.

[0056] Specifically, considering that the trained component recognition model (i.e., the subsequent target component recognition model) needs to recognize the type of components, the annotation here at least includes the annotation for the type of components. Of course, in order to improve the accuracy of subsequent model training, the annotation here should be accurate and as detailed as possible. For example, the annotation here can also include the annotation of the position of the component, etc.

[0057] In the embodiment of the present application, for ease of understanding, each piece of annotated multimodal information is stored in one data set.

[0058] Optionally, for the convenience of query or use, each piece of annotated multimodal information is structured and stored in a data set to obtain a structured data set, wherein the structured data set can be represented in the form of a table, a tree diagram, or a graphical structure.

[0059] S13, generating at least one instruction template, where the instruction template is used to describe the task that the component recognition model needs to complete.

[0060] Among them, an instruction template here can be: identifying parts in images under liquid (such as underwater), or it can be: identifying parts of mechanical equipment in images. Of course, the instruction template can also be other content, it only needs to include a description of the tasks that the part recognition model needs to complete, and there is no limitation here.

[0061] S14, constructing an instruction data training set according to each of the above instruction templates and the above data set.

[0062] Specifically, an instruction template and at least one annotated multimodal information in a data set (i.e., a multimodal information and an annotation corresponding to the multimodal information) are combined to obtain an instruction data point, and each instruction data point constitutes the above-mentioned instruction data training set.

[0063] It should be noted that the instruction templates corresponding to different instruction data points may be the same or different. For example, for different images in liquid, the corresponding instruction templates may all be "identify components in the image under liquid".

[0064] S15, training the component recognition model to be trained according to the above instruction data training set to obtain a target component recognition model.

[0065] Specifically, considering that the instruction data training set includes multimodal information of at least two modes of parts of the nuclear power plant, the above-mentioned part recognition model to be trained can select a neural network model suitable for multimodal recognition, for example, NFNet-F6 can be selected, which is a version of the NFNet (Normalizer-Free ResNets) model proposed by DeepMind, which is a deep learning model that does not require normalization. In addition, you can also choose Vision Transformer (ViT), Visual Encoder (Contrastive Language-Image Pre-trainingViT, CLIPViT) or Zhiyuan EVA-CLIP Visual Transformer (ie Eva-CLIP ViT), which is a series of models designed to significantly improve the efficiency and effectiveness of CLIP (Contrastive Language-Image Pre-training) model training.

[0066] After selecting the component recognition model to be trained, the component recognition model to be trained is trained according to the instruction data set, such as measuring the difference between the predicted value and the actual value of the component recognition model to be trained through a loss function. When the difference does not meet the stop training condition, adjust the parameters of the component recognition model until the difference meets the stop training condition.

[0067] In the embodiment of the present application, an instruction data training set is constructed according to each instruction template and data set, and then the component recognition model to be trained is trained according to the instruction data training set to obtain a target component recognition model. Since the data set is constructed according to multimodal information that at least marks the types of components of the nuclear power plant, the instruction data training set also includes multimodal information that marks the types of components of the nuclear power plant, and compared with the single modal information, the multimodal information includes more contextual information of the nuclear power plant, and the multimodal information includes images in the liquid, so that when the component recognition model to be trained is trained according to the instruction data training set, it is beneficial to improve the accuracy of the target component recognition model obtained by training to identify the types of components in the liquid of the nuclear power plant. In addition, since the instruction template is used to describe the tasks that the component recognition model needs to complete, when the component recognition model to be trained is trained according to the instruction data training set, it is equivalent to training the ability of the component recognition model to be trained to recognize tasks, which is beneficial to improve the accuracy of the target component recognition model in identifying according to the task.

[0068] In the above description, the multimodal information of the components of the nuclear power plant includes images in liquid. When acquiring the images in liquid, the following steps are included:

[0069] For components of the same nuclear power plant in liquid, images of the components are acquired when at least one of the following liquid parameters is different to obtain corresponding images in the liquid, wherein the liquid parameters include: turbidity, light, depth, angle, and background of the liquid.

[0070] The turbidity of a liquid is also referred to as the transparency of a liquid. When the liquid is water, the turbidity of the liquid refers to the turbidity of the water. Specifically, when the liquid parameter includes the turbidity of the liquid, the components can be placed in liquids of different turbidities in the pipeline, and then images of the components in liquids of different turbidities can be obtained by underwater cameras or sonar equipment to obtain corresponding images in the liquid. Optionally, in order to improve the accuracy of the target component recognition model obtained by subsequent training, the above-mentioned liquid can be set to be the same as the liquid in the pipeline of the nuclear power plant, for example, the composition and proportion of the liquid used to place the components are the same as the composition and proportion of the liquid actually existing in the pipeline of the nuclear power plant.

[0071] Among them, light refers to the light intensity or brightness in the liquid. When the liquid parameter includes light, the parts can be placed in liquids with different light in the pipeline, and then the images of the parts in the liquids with different light can be obtained through underwater cameras or sonar equipment to obtain different images in the liquid. When the liquid parameters include the turbidity of the liquid and light, the images in the liquid corresponding to the parts placed in liquids with different light and different turbidity can be obtained respectively. For example, for each light, the images in the liquid corresponding to the parts in liquids with different turbidity are obtained.

[0072] Among them, depth refers to the depth in the liquid. When the liquid parameters include depth, the components can be placed in liquids with different depths in the pipeline, and then images of the components in liquids at different depths can be obtained through underwater cameras or sonar equipment to obtain different images in the liquid.

[0073] The angle refers to the angle of the underwater camera or sonar device relative to the component placed in the liquid. When the liquid parameter includes the angle, the component can be placed in the liquid in the pipe, and then the image of the component in the liquid is obtained at different angles to obtain different images in the liquid.

[0074] The background refers to other objects in the liquid image that are not components. For example, if the acquired liquid image includes components and the internal structure of a pipe, the components belong to the foreground of the liquid image, while the internal structure of the pipe belongs to the background of the liquid image. Specifically, different liquid images are obtained by combining components with different backgrounds in the liquid and then acquiring images corresponding to the combinations.

[0075] Since the liquid parameters include at least one of the following: turbidity, light, depth, angle, and background of the liquid, when acquiring an image in the liquid based on the liquid parameters, the probability that the acquired image in the liquid contains different information can be increased, which is beneficial to improving the accuracy of the target component recognition model obtained by subsequent training in identifying components in the liquid.

[0076] In some embodiments, considering that there may be no liquid in the pipeline, in order to improve the accuracy of the trained target component recognition model in recognizing components that are not in the liquid, the multimodal information used to train the target component recognition model is set to also include a standard image. In this case, the multimodal information of the components of the nuclear power plant is obtained, including:

[0077] An image of the above-mentioned parts not in the liquid is acquired to obtain the above-mentioned standard image.

[0078] Specifically, the standard image is an image of the component not in liquid. For example, the standard image may be an image of the component in gas, such as in air.

[0079] Optionally, the step of obtaining the image of the component not in the liquid includes:

[0080] For components of the same nuclear power plant in the gas, images of the components are acquired when at least one of the following aerial parameters is different, wherein the aerial parameters include: light, angle, and background.

[0081] The light here refers to the light intensity or brightness in the gas. When the air parameters include light, the parts can be placed in the gas with different light in the pipeline, and then the image of the parts in the gas with different light can be obtained by the camera to obtain different standard images.

[0082] The angle here refers to the relative angle between the camera and the component, also known as the shooting angle. Specifically, different standard images are obtained by shooting the component at different shooting angles.

[0083] The background here refers to other objects that are not components in the standard image. Specifically, different standard images are obtained by combining components with different backgrounds in the gas and then acquiring images corresponding to the combinations.

[0084] Since the aerial parameters include at least one of the following: light, angle, and background, when a standard image is acquired based on the aerial parameters, the probability that the acquired standard image contains different information can be increased, which is beneficial to improving the accuracy of the target component recognition model obtained by subsequent training in identifying components in the gas.

[0085] In the above description, the annotation of multimodal information includes the annotation of the type of parts. In some embodiments, in order to further improve the diversity of the recognition information of the target part recognition model obtained by training, the multimodal information can be annotated in other dimensions. At this time, in the above S12, before constructing the data set according to the annotated multimodal information, it also includes:

[0086] The type of the multimodal information and at least one of the following basic information are marked to obtain the marked multimodal information, wherein the at least one basic information includes: position, state, and size.

[0087] The position here refers to the relative position of the component and other items, for example, the relative position of the component and the pipeline. For example, the position here may be: the component is on the left side of the pipeline.

[0088] Among them, the states here include: whether the component is stuck in the pipeline, whether the component blocks the pipeline, and whether the current position of the component has shifted. For example, in the normal state, the angle between the position of the component and the pipeline should be 90°, but if the current angle between the position of the component and the pipeline is less than 90° (such as 60°), the state of the component will indicate that the angle between the position of the component and the pipeline is abnormal, which is 60°.

[0089] Among them, the dimension here refers to the dimension of the component, which includes the length, width, and height of the component, and may also include the surface area, volume, etc. When the component is circular, the dimension of the component includes the diameter or radius of the component, etc.

[0090] In the embodiments of the present application, since at least one of the types of multimodal information and the information of position, state, and dimension is labeled, when training the component recognition model to be trained with the instruction data training set containing the labeled multimodal information, the component recognition model to be trained can learn more information according to the labeling result, thereby increasing the probability that the obtained target component recognition model can recognize more information.

[0091] In some embodiments, to improve the accuracy of the recognition result of the obtained target component recognition model, after training the component recognition model to be trained with the instruction data training set, the trained component recognition model is further verified with a preset instruction data verification set. At this time, the above S15, training the component recognition model to be trained according to the above instruction data training set to obtain the target component recognition model, includes:

[0092] A1. Training the component recognition model to be trained according to the above instruction data training set to obtain the first component recognition model.

[0093] Training the component recognition model to be trained according to the instruction data set, such as measuring the difference between the predicted value and the actual value of the component recognition model to be trained through a loss function. If the difference does not meet the stop training condition, adjusting the parameters of the component recognition model until the difference meets the stop training condition, and the obtained model is the above first component recognition model.

[0094] A2. Verifying the above first component recognition model with a preset instruction data verification set, where the above instruction data verification set is constructed according to each of the above instruction templates and a preset verification set, and the samples included in the above verification set are similar to the samples included in the above data set.

[0095] Among them, the samples included in the preset verification set are different from those included in the data set, but the sample distribution in the verification set is similar to that in the data set. For example, assuming that in the data set, the distribution ratios of images in the liquid in terms of turbidity, light, depth, angle, and background of the liquid are: 1:2:2:3:2, then in the verification set, the distribution ratios of images in the liquid in terms of turbidity, light, depth, angle, and background of the liquid are also: 1:2:2:3:2.

[0096] After determining the verification set, an instruction data verification set is constructed according to the above instruction template and the verification set. The instruction data points in the instruction data verification set are similar to those in the instruction data training set. By setting it in this way, it is possible to better verify the accuracy of the recognition result of the first component recognition model trained according to the instruction data training set for the instruction data verification set with similar instruction data points.

[0097] In the embodiment of the present application, a preset instruction data verification set is used to verify the first component recognition model, that is, the first component recognition model is used to recognize the instruction data points in the instruction data verification set. After obtaining the recognition result, the recognition result is compared with the annotation corresponding to the instruction data point, and the verification result is obtained according to the comparison result.

[0098] A3. Adjust the parameters of the above instruction template and / or the above first component recognition model according to the verification result to obtain the above target component recognition model.

[0099] Specifically, when the verification result does not meet the user's requirements, for example, when the verification result indicates that the accuracy of the recognition result of the first component recognition model is low, it is possible to choose to adjust the instruction template, or choose to adjust the parameters of the first component recognition model (such as learning rate, number of iterations, etc.), or choose to adjust both the instruction template and the parameters of the first component recognition model.

[0100] In the embodiment of the present application, after training the component recognition model to be trained according to the instruction data training set, a preset instruction data verification set is also used to verify the trained first component recognition model. Since the instruction data verification set is constructed according to each instruction template and the preset verification set, and the samples included in the verification set are similar to those included in the data set, the samples included in the instruction data verification set are also similar to those included in the instruction data training set. That is, the verification result obtained by verifying the first component recognition model according to the instruction data verification set can accurately indicate whether the first component recognition model meets the training requirements (or expected results). Therefore, after adjusting the parameters of the instruction template and / or the first component recognition model according to the verification result, it is beneficial to improve the accuracy of the obtained target component recognition model.

[0101] In the above description, the parameters of the instruction template and / or the first component recognition model may be adjusted according to the verification result. In some embodiments, considering that adjusting the instruction template is simpler than adjusting the parameters of the first component recognition model, the instruction template may be adjusted first. In this case, the above A3 includes:

[0102] A31. If the above verification result does not meet the expected result, adjust the above instruction template.

[0103] Specifically, when the verification result indicates that the recognition result of the instruction data point in the instruction verification data set by the first component recognition model does not meet the expected result (such as the accuracy of the recognition result is lower than the preset accuracy threshold), it is selected to adjust the instruction template first, for example, adjust the order of words in the instruction template; for example, delete words without specific meaning in the instruction template; for example, add directional (such as left, right) words or range words (such as from... to..., in..., within...), etc. in the instruction template.

[0104] Optionally, the above A31 includes:

[0105] When the above verification results do not meet the expected results, detect whether the samples in the above verification set meet the sample requirements; when the samples in the above verification set meet the sample requirements, adjust the above instruction template.

[0106] Among them, the sample requirements here include: accurate labeling and the distribution ratio of different types of samples meeting the requirements.

[0107] In the embodiment of the present application, if it is determined that the verification result does not meet the expected result, it is first determined whether the samples in the verification set meet the sample requirements. If the sample requirements are not met, it indicates that the instruction data verification set constructed based on the verification set also does not meet the sample requirements. At this time, it is necessary to adjust the samples in the verification set, and then generate an adjusted instruction data verification set based on the adjusted verification set. Finally, the first component recognition model is verified based on the adjusted instruction data verification set, and the obtained verification result is compared with the expected result. If the samples in the verification set meet the sample requirements, the instruction template is adjusted. The above processing is conducive to improving the accuracy of the timing of adjusting the instruction template.

[0108] A32. Generate a new instruction data verification set based on the adjusted instruction template.

[0109] Specifically, the adjusted instruction template and the preset verification set are constructed to obtain a new instruction data verification set. The new instruction data verification set includes multiple new instruction data points, and the new instruction data points are determined in the following manner: an adjusted instruction template and at least one annotated multimodal information in the verification set (i.e., a multimodal information and an annotation corresponding to the multimodal information) are combined to obtain a new instruction data point in the new instruction data verification set.

[0110] A33. Use the new instruction data verification set to verify the first component recognition model to obtain a new verification result.

[0111] Among them, the process of using the new instruction data verification set to verify the first component recognition model is similar to the process of using the instruction data verification set to verify the first component recognition model, which will not be repeated here.

[0112] A34. When the new verification result does not meet the expected result, adjust the parameters of the first component identification model to obtain the target component identification model.

[0113] Among them, the number of times the parameters of the first component identification model are adjusted can be one or more than one. For example, when the new verification result does not meet the expected result, the parameters of the first component identification model are adjusted, and then the first component identification model after the parameter adjustment is verified using a new instruction data verification set to determine whether the obtained verification result meets the expected result. If it still does not meet the expected result, continue to return to the step of adjusting the parameters of the first component identification model and subsequent steps until it is determined that the obtained verification result meets the expected result. The first component identification model after the parameter adjustment is the above-mentioned target component identification model. Of course, if the new verification result meets the above-mentioned expected result, there is no need to adjust the parameters of the first component identification model, and the first component identification model is the above-mentioned target component identification model.

[0114] In the embodiment of the present application, when it is determined that the verification result does not meet the expected result, the instruction template is adjusted first, and then the parameters of the first component identification model are adjusted after it is determined that the new verification result still does not meet the expected result. Since the parameters of the first component identification model will affect all instruction data points after being adjusted, and the adjustment of a single instruction template will only affect itself, it is simpler to adjust the instruction template than to adjust the parameters of the first component identification model. That is, the above adjustment method is conducive to simplifying the complexity of obtaining the target component identification model.

[0115] In some embodiments, considering that instruction fine-tuning is intended to specialize the pre-trained model into a model for a specific task, in order to further improve the accuracy of the obtained target component recognition model in identifying the components of a nuclear power plant, the parameters of the target component recognition model may be fine-tuned. For example, by using an efficient parameter fine-tuning method such as low-rank adaptation (LoRA), a small amount of parameters of the target component recognition model may be fine-tuned to reduce the training cost and maintain the performance of the target component recognition model.

[0116] The above describes the process of obtaining the target component recognition model based on the multimodal information training of the components of the nuclear power plant. Since these multimodal information are not fused during the training process, the target component recognition model obtained according to the above method is equivalent to that obtained by training based on the single-modal information of multiple modes. Since the amount of information contained in the single-modal information of different modes after fusion is greater, and the use of the fused information to train the model is conducive to improving the generalization of the model obtained after training, the fused single-modal information of different modes can be used to train the component recognition model to be trained. At this time, when the above A34 adjusts the parameters of the above first component recognition model to obtain the above target component recognition model, it includes:

[0117] B1. Adjust the parameters of the first component identification model to obtain a second component identification model.

[0118] The adjustment here is the same as the above-mentioned step A34, except that after the parameters of the first component identification model are adjusted, the model obtained is the second component identification model.

[0119] B2. For the modal information of each mode of the above-mentioned component in the above-mentioned data set, extract features corresponding to the above-mentioned modal information.

[0120] For example, when the data set includes modal information corresponding to two modalities, image and audio information, the features corresponding to the modal information of the image modality are extracted, and the features corresponding to the modal information of the audio modality are extracted. Of course, if there are multiple modal information of the same modality, the features of each modal information of the modality need to be extracted.

[0121] For example, when the dataset includes modal information corresponding to the two modalities of image and text information:

[0122] For the image modality, features with specific semantic information can be extracted from the image to obtain a set of image features. These image features can be global features of the image (such as color histogram, texture features, etc.) or local features of the image (such as corners, edges, regions, etc.). By extracting image features, the pixel information in the image is converted into a high-level semantic representation that is useful for subsequent tasks (such as classification, detection, recognition, etc.).

[0123] For text information, features with specific semantic information can be extracted from text information through natural language processing (NLP), word embedding models Word2Vec, BERT, etc. to obtain a set of text features. These text features can be language units such as words, phrases, sentences, etc. in the text, or semantic relations and contextual information between them. By extracting text features, the language information in the text is converted into a high-level semantic representation that is useful for subsequent tasks (such as text results returned after underwater image recognition, etc.).

[0124] B3. Fusing the features of each modality to obtain multiple fused features, wherein each of the fused features includes features of at least two modalities.

[0125] Specifically, considering that the multimodal information at least includes the modal information corresponding to the image in the liquid, and the modal information corresponding to at least one of the text information, video information and audio information, at least two modal features can be extracted from the multimodal information of the parts. When the multimodal information has modal information of only two modalities, the features of the two modalities extracted from the modal information of the two modalities can be fused to obtain a fused feature. It should be pointed out that since there may be multiple modal information of one modality, for example, for the modal information of the image type, there is modal information corresponding to the image in the liquid, and there is also modal information corresponding to the standard image. Therefore, the number of features extracted from the modal information of the two modalities may be more than 2, that is, the number of features corresponding to a fused feature may also be more than 2. Of course, if only one feature is extracted from one modal information of each modality for fusion, the number of features corresponding to a fused feature is equal to 2.

[0126] Optionally, when the number of modalities corresponding to the multimodal information is greater than 2, the number of modalities to be fused can be selected from 2 to M (M is the number of modalities), and the corresponding modality is selected according to the number of modalities to be fused, and then the corresponding features are extracted from the modal information of the selected modality, and finally the extracted features are fused to obtain the fused features.

[0127] In an embodiment of the present application, the features of each mode extracted can be encoded into a vector representation that can capture the key information in each mode, and then fused. Since feature encoding can convert raw data into a more abstract representation, when the fused features obtained by fusion of the feature-encoded vectors are used for training, it helps the trained model to capture the inherent structure and model of the raw data, thereby facilitating the improvement of the generalization ability of the trained target component recognition model. For example, when the extracted features are image features and text features, the extracted image features and text features can be first encoded into vector representations of fixed lengths, respectively. These vector representations should be able to capture the key information in each mode and have a certain generalization ability, and then the two vector representations obtained by encoding are fused to obtain fused features.

[0128] B4. Generate a fusion instruction, and use the fusion instruction and the multiple fusion features to train the second component recognition model to obtain the target component recognition model.

[0129] The number of the above fusion instructions is greater than or equal to 1. The fusion instruction is similar to the instruction template, and is also used to describe the task that the component recognition model needs to complete. For example, the fusion instruction can be: recognize matching components in images and texts.

[0130] Specifically, the fusion instruction is combined with the fusion feature, and then the second component recognition model is trained according to the obtained multiple combination results. The combination process is similar to the process of generating instruction data points, which will not be repeated here.

[0131] In the embodiment of the present application, after adjusting the parameters of the first component recognition model to obtain the second component recognition model, the corresponding features are extracted from the modal information of each mode of the component, and the features of different modes are fused to obtain fused features, and then the second component recognition model is trained according to the generated fusion instructions and fused features. Since the fused features corresponding to multiple modes contain more contextual information than the features of a single mode, the use of the fused features to train the second component recognition model is conducive to improving the generalization and robustness of the obtained target component recognition model.

[0132] In some embodiments, considering that the distances between the features of different modalities included in the fusion features of different semantic spaces are usually different, and the smaller the distance between the features of different modalities, the higher the accuracy of the recognition result of the target component recognition model obtained by training according to the fusion features, therefore, the fusion features corresponding to the semantic space that minimizes the distance between the features of different modalities can be selected as the fusion features for subsequent training. In this case, the above B3 includes:

[0133] B31. Determine at least two mapping methods, and for any of the above mapping methods, map the features of each modality to the same semantic space to obtain corresponding mapping features.

[0134] The mapping method here refers to the method of mapping the features of each modality into the same semantic space.

[0135] Optionally, the mapping method includes simple concatenation and mapping through a mapping function or a mapping matrix. The mapping function may be a function corresponding to weighted summation, that is, corresponding weights are set for features of different modes, the weights and features of the same mode are multiplied, and then added to the product of other modes (that is, the product obtained by multiplying the weights and features).

[0136] Optionally, the weights of different modalities can be determined by a self-attention mechanism or a graph neural network. In this way, a greater weight is given to the features of the modality of interest to increase the proportion of the feature in the fusion feature.

[0137] In the embodiment of the present application, mapping the features of each modality to the same semantic space is equivalent to aligning the features of different modalities in space or time (ie, modality alignment), so that they can be compared and associated in the same framework.

[0138] B32. For any semantic space, calculate the differences between different features in the above-mentioned mapping features in the semantic space.

[0139] In the embodiment of the present application, for each mapping feature in any semantic space, the difference between different features in the mapping feature can be calculated. Specifically, the difference between different features in the mapping feature can be calculated by methods such as cosine similarity, Euclidean distance, or Manhattan distance.

[0140] B33. In each semantic space, determine the semantic space corresponding to the minimum value of the difference between different features in the same mapping feature to obtain the target semantic space.

[0141] Specifically, for any mapping feature, the minimum value of the difference can be selected from each semantic space, and then a minimum value can be selected from the minimum values ​​of the differences in all mapping features. The semantic space corresponding to the finally selected minimum value is the target semantic space mentioned above.

[0142] For example, suppose there are three mapping features: mapping feature 1, mapping feature 2, and mapping feature 3, and suppose there are two semantic spaces, semantic space 1 and semantic space 2. In semantic space 1, the difference between different features in mapping feature 1 is difference 11, the difference between different features in mapping feature 2 is difference 12, and the difference between different features in mapping feature 3 is difference 13. In semantic space 2, the difference between different features in mapping feature 1 is difference 21, the difference between different features in mapping feature 2 is difference 22, and the difference between different features in mapping feature 3 is difference 23. Assuming that difference 11 is less than difference 21, difference 12 is less than difference 22, and difference 23 is less than difference 13, first select difference 11, difference 12, and difference 23, and then select the minimum value among difference 11, difference 12, and difference 23, assuming it is difference 23, then the semantic space corresponding to difference 23 is the target semantic space, that is, semantic space 2 is the above-mentioned target semantic space.

[0143] B34. Use the mapping features in the target semantic space as the above fusion features.

[0144] Specifically, each mapping feature in the target semantic space is used as the multiple fusion features mentioned above.

[0145] In the embodiment of the present application, the features of each modality are mapped to the corresponding semantic space by at least two different mapping methods to obtain the corresponding mapping features. Since the distance (or difference) between different features in the same mapping feature is usually different in different semantic spaces, the mapping feature with the minimum distance can be selected from different semantic spaces to determine the fusion feature for training. Since the smaller the distance between the features of the corresponding different modalities in the fusion feature, the more accurate the recognition result obtained by the target component recognition model trained by the fusion feature, the above method is conducive to improving the accuracy of the target component recognition model obtained subsequently. In addition, since the component recognition model training method provided by the embodiment of the present application is to identify foreign matter in liquid (such as underwater), and the recognition process may involve information of multiple modalities, such as optical images, depth information, etc. These modal information sources are different, and the content of expression is also different. Through modal alignment, the target component recognition model can understand the meaning of different modal information, and associate them with the actual features of underwater foreign matter, and establish a connection between different modalities, such as corresponding a certain area in the optical endoscope image to numerical values ​​such as water depth. That is, through modal alignment, it is helpful to increase the probability of the target component recognition model learning a richer feature representation, thereby improving its generalization ability in the underwater foreign body image recognition task. On the other hand, when performing feature fusion of different modalities, information from different modalities can be integrated together to form a more comprehensive and accurate description of underwater foreign objects. This integration helps the target component recognition model capture the complementarity between different modalities, thereby providing more reliable recognition results, improving recognition precision and accuracy, and reducing dependence on single modal information, enhancing its robustness in complex underwater environments.

[0146] The above introduces the process of obtaining a target component recognition model through an instruction data training set, or through an instruction data training set and an instruction data verification set. Before applying the target component recognition model to recognize images in liquids, in order to further improve the accuracy of the target component recognition model in recognizing components, the trained target component recognition model can be tested.

[0147] Specifically, a test set is first determined, and then the target component recognition model is tested using the test set.

[0148] In order to achieve a more comprehensive detection of the target parts recognition model, it is necessary to construct a more reasonable test set. For example, the samples in the constructed test set are different from the samples in the data set (or validation set) and are not completely similar.

[0149] When constructing a test set, the acquired test samples, such as images (which may be photographed images or images extracted from videos), may be subjected to at least one transformation, such as translation, rotation, scaling, and adding noise, to generate more test samples. This process helps to evaluate the robustness of the target component recognition model when processing blurred, deformed, or noise-interfered images.

[0150] Optionally, in addition to increasing the test samples by the above method, the test samples can also be increased by generating synthetic images. For example, using generative adversarial networks (GANs) or other synthesis techniques, synthetic images that are similar to but not completely the same as the images under real liquids are generated, such as synthesizing images of more extreme or rare scenes. In this way, when these synthetic images are used as test samples to test the target component recognition model, it is helpful to evaluate the recognition ability of the target component recognition model in these extreme scenes.

[0151] Optionally, in addition to increasing the test samples in the above manner, the test samples can also be increased by generating images containing specific details. For example, special test samples are designed for specific details of parts (such as arcs, corners, slopes, etc.). When these test samples are used to test the target part recognition model, it is helpful to evaluate the recognition ability of the target part recognition model for these details.

[0152] Optionally, in addition to increasing the test samples in the above manner, the test samples may also be increased by generating images of different scales. For example, for the same image, images of different image resolutions corresponding to the image are generated. Since the resolutions of the equipment used to collect images in the pipeline of the nuclear power plant may be different, when images of different image resolutions are used as test samples to test the target component recognition model, the purpose of optimizing the recognition ability of the details of the target component recognition model can be achieved by observing the performance differences of the target component recognition model at different scales and making adaptive adjustments.

[0153] In some embodiments, when the target component recognition model is tested using a test set, the response time (or processing speed) of the target component recognition model can also be determined. If the response time does not meet the preset response requirements, the target component recognition model is adjusted to shorten the response time of the target component recognition model.

[0154] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0155] After the target component recognition model is obtained through training in the above manner, the target component recognition model can be used to recognize images acquired by the nuclear power plant.

[0156] Figure 2 A schematic flow chart of a component identification method provided in an embodiment of the present application is shown. The component identification method is applied to a nuclear power plant, such as an electronic device in a nuclear power plant, and is described in detail as follows:

[0157] S21, obtaining an image to be recognized.

[0158] The image to be identified is an image that needs to be identified as a component, and may be an image captured in a pipeline containing liquid or gas in a nuclear power plant.

[0159] Specifically, the electronic device may use an image sent by other devices (such as a camera) as the image to be recognized, or may use an image acquired by itself as the image to be recognized, which is not limited here.

[0160] S22, using the target component recognition model described above to recognize the image to be recognized, to obtain a recognition result.

[0161] The recognition result includes: when there is a component in the image to be recognized, indicating the type of the component. Optionally, the recognition result is also used to indicate at least one of the position, state and size of the component.

[0162] In the embodiment of the present application, since the target component recognition model has a high accuracy in identifying the type of components in the liquid of the nuclear power plant according to the task, when the image to be recognized is recognized according to the target component recognition model, it is helpful to improve the accuracy of the recognition result obtained. In addition, since no human participation is required when recognizing the image to be recognized, the recognition result will not be affected by personal experience, thereby further improving the accuracy of the recognition result obtained.

[0163] In some embodiments, after the above S22, the method further includes:

[0164] When the recognition result includes the type of the parts of the nuclear power plant, the parts image corresponding to the type of the parts is determined according to the pre-stored mapping relationship, and the parts image is displayed. The pre-stored mapping relationship is used to record the correspondence between the type of parts and the parts image.

[0165] The above-mentioned component image may be a two-dimensional image or a three-dimensional image, which is not limited here.

[0166] Specifically, a mapping relationship between the type of a component and the image of the component (i.e., the component image) is pre-stored. When the recognition result obtained by the target component recognition model indicates that the image to be recognized has a component and the type (or model) of the component is given, the corresponding component image can be found from the mapping relationship and displayed. Since the component image is displayed, it is beneficial for the user to obtain the image information of the component more intuitively, thereby improving the accuracy of the evaluation of the component.

[0167] Corresponding to the component recognition model training method described in the above embodiment, Figure 3 A structural block diagram of a component recognition model training device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0168] Reference Figure 3 The component recognition model training device 3 is applied to a nuclear power plant, and includes: a multimodal information acquisition module 31, a data set construction module 32, an instruction template generation module 33, an instruction data training set construction module 34, and a target component recognition model generation module 35. Among them:

[0169] The multimodal information acquisition module 31 is used to acquire multimodal information of components of a nuclear power plant, wherein the multimodal information includes images in liquid and at least one of the following information: text information, video information and audio information, wherein the images in liquid are images of the components in liquid.

[0170] The data set construction module 32 is used to construct a data set according to the annotated multimodal information, where the annotated multimodal information includes the type of components.

[0171] The instruction template generating module 33 is used to generate at least one instruction template, and the instruction template is used to describe the task that the component recognition model needs to complete.

[0172] The instruction data training set construction module 34 is used to construct an instruction data training set according to each of the above instruction templates and the above data set.

[0173] The target component recognition model generation module 35 is used to train the component recognition model to be trained according to the above instruction data training set to obtain the target component recognition model.

[0174] In the embodiment of the present application, an instruction data training set is constructed according to each instruction template and data set, and then the component recognition model to be trained is trained according to the instruction data training set to obtain a target component recognition model. Since the data set is constructed according to multimodal information that at least marks the types of components of the nuclear power plant, the instruction data training set also includes multimodal information that marks the types of components of the nuclear power plant, and compared with the single modal information, the multimodal information includes more contextual information of the nuclear power plant, and the multimodal information includes images in the liquid, so that when the component recognition model to be trained is trained according to the instruction data training set, it is beneficial to improve the accuracy of the target component recognition model obtained by training to identify the types of components in the liquid of the nuclear power plant. In addition, since the instruction template is used to describe the tasks that the component recognition model needs to complete, when the component recognition model to be trained is trained according to the instruction data training set, it is equivalent to training the ability of the component recognition model to be trained to recognize tasks, which is beneficial to improve the accuracy of the target component recognition model in identifying according to the task.

[0175] In some embodiments, when acquiring the above-mentioned image in the liquid, it includes:

[0176] For components of the same nuclear power plant in liquid, images of the components are acquired when at least one of the following liquid parameters is different to obtain corresponding images in the liquid, wherein the liquid parameters include: turbidity, light, depth, angle, and background of the liquid.

[0177] In some embodiments, the multimodal information further includes a standard image, and the multimodal information acquisition module 31 is specifically used for:

[0178] An image of the above-mentioned parts not in the liquid is acquired to obtain the above-mentioned standard image.

[0179] In some embodiments, obtaining the image of the component not in the liquid includes:

[0180] For components of the same nuclear power plant in the gas, images of the components are acquired when at least one of the following aerial parameters is different, wherein the aerial parameters include: light, angle, and background.

[0181] In some embodiments, the component recognition model training device 3 provided in the embodiment of the present application further includes:

[0182] The labeling module is used to label the type of the above multimodal information and at least one of the following basic information to obtain the labeled multimodal information, and the at least one basic information includes: position, state, and size.

[0183] In some embodiments, the target component recognition model generation module 35 includes:

[0184] The first component recognition model generating unit is used to train the component recognition model to be trained according to the above instruction data training set to obtain a first component recognition model.

[0185] A verification unit is used to verify the above-mentioned first component recognition model using a preset instruction data verification set, wherein the above-mentioned instruction data verification set is constructed according to each of the above-mentioned instruction templates and the preset verification set, and the samples included in the above-mentioned verification set are similar to the samples included in the above-mentioned data set.

[0186] A parameter selection and adjustment unit is used to adjust the parameters of the above-mentioned instruction template and / or the above-mentioned first component recognition model according to the verification result to obtain the above-mentioned target component recognition model.

[0187] In some embodiments, the parameter selection and adjustment unit includes:

[0188] The verification result and expected result comparison unit is used to adjust the above instruction template when the above verification result does not meet the expected result.

[0189] A new instruction data verification set generating unit is used to generate a new instruction data verification set according to the adjusted instruction template.

[0190] The new verification result obtaining unit is used to verify the first component recognition model using the new instruction data verification set to obtain a new verification result.

[0191] The parameter adjustment unit of the first component identification model is used to adjust the parameters of the first component identification model to obtain the target component identification model when the new verification result does not meet the expected result.

[0192] In some embodiments, the verification result and expected result comparison unit includes:

[0193] When the above verification results do not meet the expected results, detect whether the samples in the above verification set meet the sample requirements; when the samples in the above verification set meet the sample requirements, adjust the above instruction template.

[0194] In some embodiments, when the parameter adjustment unit of the first component recognition model adjusts the parameters of the first component recognition model to obtain the target component recognition model, it is specifically used to:

[0195] The parameters of the first component recognition model are adjusted to obtain a second component recognition model; for the modal information of each mode of the component in the data set, features corresponding to the modal information are extracted; the features of each mode are fused to obtain a plurality of fused features, wherein each of the fused features includes features of at least two modes; a fusion instruction is generated, and the second component recognition model is trained using the fusion instruction and the plurality of fused features to obtain the target component recognition model.

[0196] In some embodiments, the above-mentioned fusion of features of each modality to obtain multiple fusion features includes:

[0197] Determine at least two mapping methods, and for any of the mapping methods, map the features of each modality to the same semantic space to obtain corresponding mapping features;

[0198] For any of the above semantic spaces, calculating the differences between different features in the above mapping features in the above semantic space;

[0199] In each of the above semantic spaces, determine the semantic space corresponding to the minimum value of the difference between different features in the same mapping feature to obtain the target semantic space;

[0200] The mapping features in the above target semantic space are used as the above fusion features.

[0201] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0202] Corresponding to the component identification method described in the above embodiment, Figure 4 A structural block diagram of a component identification device provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.

[0203] Reference Figure 4 The component recognition device 4 is applied to a nuclear power plant and includes a to-be-recognized image acquisition module 41 and a recognition result generation module 42. Among them:

[0204] The to-be-recognized image acquisition module 41 is used to acquire the to-be-recognized image.

[0205] The recognition result generating module 42 is used to recognize the above-mentioned image to be recognized by using the target component recognition model described above to obtain a recognition result.

[0206] In the embodiment of the present application, since the target component recognition model has a high accuracy in identifying the type of components in the liquid of the nuclear power plant according to the task, when the image to be recognized is recognized according to the target component recognition model, it is helpful to improve the accuracy of the recognition result obtained. In addition, since no human participation is required when recognizing the image to be recognized, the recognition result will not be affected by personal experience, thereby further improving the accuracy of the recognition result obtained.

[0207] In some embodiments, the component identification device 4 provided in the embodiment of the present application further includes:

[0208] The component image determination module is used to determine the component image corresponding to the type of the component according to the pre-stored mapping relationship after the above-mentioned recognition result is obtained, when the above-mentioned recognition result includes the type of the component of the nuclear power plant, wherein the above-mentioned pre-stored mapping relationship is used to record the correspondence between the type of the component and the component image.

[0209] The component image display module is used to display the above component images.

[0210] The above-mentioned component image may be a two-dimensional image or a three-dimensional image, which is not limited here.

[0211] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0212] Figure 5 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Figure 5 As shown, the electronic device 5 of this embodiment includes: at least one processor 50 ( Figure 5 Only one processor is shown in the figure), a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, wherein the processor 50 implements the steps in any of the above-mentioned method embodiments when executing the computer program 52.

[0213] The electronic device 5 may be a computing device such as a desktop computer, a notebook, a PDA, or a cloud server. The electronic device may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will appreciate that Figure 5 It is only an example of the electronic device 5 and does not constitute a limitation on the electronic device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0214] The processor 50 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0215] In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. In other embodiments, the memory 51 may also be an external storage device of the electronic device 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 5. Further, the memory 51 may also include both an internal storage unit and an external storage device of the electronic device 5. The memory 51 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory 51 may also be used to temporarily store data that has been output or is to be output.

[0216] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0217] An embodiment of the present application also provides a network device, which includes: at least one processor, a memory, and a computer program stored in the memory and executable on the at least one processor, wherein the processor implements the steps in any of the above-mentioned method embodiments when executing the computer program.

[0218] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.

[0219] An embodiment of the present application provides a computer program product. When the computer program product runs on an electronic device, the electronic device can implement the steps in the above-mentioned method embodiments when executing the computer program product.

[0220] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device that can carry the computer program code to the camera / electronic device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, RandomAccess Memory), electric carrier signal, telecommunication signal and software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electric carrier signals and telecommunication signals.

[0221] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0222] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0223] In the embodiments provided in the present application, it should be understood that the disclosed devices / network equipment and methods can be implemented in other ways. For example, the device / network equipment embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0224] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0225] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A component recognition model training method, characterized in that: Applications in nuclear power plants, including: Acquire multimodal information of components of a nuclear power plant, the multimodal information comprising an image in liquid and at least one of the following information: text information, video information, and audio information, wherein the image in liquid is an image of the component in liquid; Constructing a data set according to the annotated multimodal information, wherein the annotated multimodal information includes the type of the component; Generate at least one instruction template, where the instruction template is used to describe a task that needs to be completed by the component recognition model; Constructing an instruction data training set according to each of the instruction templates and the data set; The component recognition model to be trained is trained according to the instruction data training set to obtain a target component recognition model.

2. The component recognition model training method according to claim 1, characterized in that: When acquiring the image in the liquid, the method includes: For components of the same nuclear power plant in liquid, images of the components under different at least one of the following liquid parameters are acquired to obtain corresponding images in the liquid, wherein the liquid parameters include: turbidity, light, depth, angle, and background of the liquid.

3. The component recognition model training method according to claim 1, characterized in that: The multimodal information also includes a standard image, and the multimodal information of the components of the nuclear power plant is obtained, including: An image of the component not in the liquid is acquired to obtain the standard image.

4. The component recognition model training method according to claim 3, characterized in that: The step of obtaining an image of the component that is not in the liquid comprises: For components of the same nuclear power plant in the gas, images of the components are acquired when at least one of the following aerial parameters is different, wherein the aerial parameters include: light, angle, and background.

5. The component recognition model training method according to any one of claims 1 to 4, characterized in that: Before constructing a data set according to the annotated multimodal information, the method further includes: The type of the multimodal information and at least one of the following basic information are labeled to obtain the labeled multimodal information, where the at least one basic information includes: position, state, and size.

6. The component recognition model training method according to claim 5, characterized in that: The training of the component recognition model to be trained according to the instruction data training set to obtain the target component recognition model comprises: Training the component recognition model to be trained according to the instruction data training set to obtain a first component recognition model; Using a preset instruction data verification set to verify the first component recognition model, wherein the instruction data verification set is constructed according to each of the instruction templates and a preset verification set, and the samples included in the verification set are similar to the samples included in the data set; The parameters of the instruction template and / or the first component recognition model are adjusted according to the verification result to obtain the target component recognition model.

7. The component recognition model training method according to claim 6, characterized in that: The step of adjusting the parameters of the instruction template and / or the first component recognition model according to the verification result to obtain the target component recognition model includes: If the verification result does not meet the expected result, adjusting the instruction template; Generate a new instruction data verification set according to the adjusted instruction template; Using the new instruction data verification set to verify the first component recognition model to obtain a new verification result; When the new verification result does not meet the expected result, the parameters of the first component recognition model are adjusted to obtain the target component recognition model.

8. The component recognition model training method according to claim 7, characterized in that: When the verification result does not meet the expected result, adjusting the instruction template includes: If the verification result does not meet the expected result, detecting whether the samples in the verification set meet the sample requirements; When the samples in the verification set meet the sample requirements, the instruction template is adjusted.

9. The component recognition model training method according to claim 7, characterized in that: The step of adjusting the parameters of the first component recognition model to obtain the target component recognition model includes: Adjusting the parameters of the first component identification model to obtain a second component identification model; For each mode of the modal information of the component in the data set, extract features corresponding to the modal information; Fusing the features of each modality to obtain a plurality of fused features, wherein each of the fused features includes features of at least two modalities; A fusion instruction is generated, and the second component recognition model is trained using the fusion instruction and the multiple fusion features to obtain the target component recognition model.

10. The component recognition model training method according to claim 9, characterized in that: The features of each modality are fused to obtain multiple fused features, including: Determine at least two mapping methods, and for any of the mapping methods, map the features of each modality to the same semantic space to obtain corresponding mapping features; For any of the semantic spaces, calculating differences between different features in the mapping features in the semantic space; In each of the semantic spaces, determining the semantic space corresponding to the minimum value of the difference between different features in the same mapping feature, to obtain a target semantic space; The mapping features in the target semantic space are used as the fusion features.

11. A component identification method, characterized in that: Applications in nuclear power plants, including: Obtain an image to be recognized; The target component recognition model according to any one of claims 1 to 10 is used to recognize the image to be recognized to obtain a recognition result.

12. The component identification method according to claim 11, characterized in that: After obtaining the recognition result, the method further includes: In the case where the recognition result includes the type of a component of a nuclear power plant, determining a component image corresponding to the type of the component according to a pre-stored mapping relationship, wherein the pre-stored mapping relationship is used to record the correspondence between the type of the component and the component image; An image of the component is displayed.

13. A component recognition model training device, characterized in that: Applications in nuclear power plants, including: A multimodal information acquisition module, used to acquire multimodal information of components of a nuclear power plant, wherein the multimodal information includes an image in liquid and at least one of the following information: text information, video information, and audio information, wherein the image in liquid is an image of the component in liquid; A data set construction module, used to construct a data set according to the annotated multimodal information, wherein the annotated multimodal information includes the type of the component; An instruction template generation module, used to generate at least one instruction template, wherein the instruction template is used to describe the task that the component recognition model needs to complete; An instruction data training set construction module, used to construct an instruction data training set according to each of the instruction templates and the data set; The target component recognition model generation module is used to train the component recognition model to be trained according to the instruction data training set to obtain the target component recognition model.

14. A component identification device, characterized in that: Applications in nuclear power plants, including: An image acquisition module to be identified, used to acquire an image to be identified; The recognition result generating module is used to recognize the image to be recognized by using the target component recognition model according to any one of claims 1 to 10 to obtain a recognition result.

15. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method according to any one of claims 1 to 10 is implemented, or the method according to claim 11 or 12 is implemented.

16. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 10 is implemented, or the method according to claim 11 or 12 is implemented.

17. A computer program product, characterized in that The method comprises a computer program, which, when executed, implements the method according to any one of claims 1 to 10, or implements the method according to claim 11 or 12.