Image recognition method, device, electronic device, and computer-readable storage medium
By performing device area detection and local feature fusion on the client, the problem of low accuracy of component recognition in complex devices is solved, and fast and accurate part recognition is achieved.
Patent Information
- Application Number
- CN202110870207.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-30
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-07-30
AI Technical Summary
The prior art has low accuracy when identifying parts in complex equipment, especially due to the difficulty of identification caused by the tiny or obscured parts.
By displaying the part recognition page on the client, the device area detection is performed, the area selection page is displayed when the threshold exceeds the threshold, and the part recognition results are displayed in response to the user's selection, combining local feature extraction and fusion to identify the device area and type.
Improve the accuracy of image recognition and achieve fast and accurate part recognition.
Smart Images

Figure CN113822295B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of communication technologies, and in particular to an image recognition method, device, electronic device, and computer-readable storage medium. Background Art
[0002] In recent years, with the rapid development of internet technology, the application of image recognition has become increasingly widespread. For example, it can identify the parts contained in a device image. In the field of parts identification, existing image recognition methods often directly identify part information in device images.
[0003] During the research and practice of the prior art, the inventors of the present invention discovered that when directly identifying component information in a device image, complex devices often contain many components, and some components are very small, or even blocked by some spare parts, resulting in a low accuracy rate in component identification, thereby reducing the accuracy of image recognition. Summary of the Invention
[0004] Embodiments of the present invention provide an image recognition method, device, electronic device, and computer-readable storage medium, which can improve the accuracy of image recognition.
[0005] An image recognition method, comprising:
[0006] Displaying a parts identification page of a client, wherein the parts identification page includes an image acquisition control;
[0007] In response to a triggering operation on the image acquisition control, performing device area detection on the acquired target device image;
[0008] When the number of detected device areas exceeds a preset number threshold, a region selection page is displayed, wherein the region selection page includes multiple device areas to be selected;
[0009] In response to a selection operation on the device area, a parts identification result page is displayed, the parts identification result page including a parts list of the target device area selected by the selection operation, the parts list including attribute information of at least one device part in the target device area.
[0010] Optionally, an embodiment of the present invention further provides another image recognition method, including:
[0011] receiving a target device image sent by a terminal, and performing local feature extraction on the target device image to obtain image features of multiple sizes;
[0012] fusing the image features to obtain fused image features of the target device image;
[0013] identifying at least one device region containing device parts in the target device image based on the fused image features;
[0014] Based on the fused image features, the device area is classified, and device part information corresponding to the type of the device area is obtained to obtain device area identification information of the target device image;
[0015] The device area identification information is sent to the terminal, so that the terminal performs device area detection on the target device image based on the device area identification information.
[0016] Accordingly, an embodiment of the present invention provides an image recognition device, comprising:
[0017] A first display unit is used to display a part identification page of a client, wherein the part identification page includes an image acquisition control;
[0018] a detection unit, configured to perform device area detection on the acquired target device image in response to a triggering operation on the image acquisition control;
[0019] A second display unit is configured to display an area selection page when the number of detected device areas exceeds a preset number threshold, the area selection page including a plurality of device areas to be selected;
[0020] The third display unit is used to display a part identification result page in response to a selection operation on the device area, wherein the part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area.
[0021] Optionally, an embodiment of the present invention further provides another image recognition device, including:
[0022] a receiving unit, configured to receive a target device image sent by a terminal, and extract local features of the target device image to obtain image features of multiple sizes;
[0023] a fusion unit, configured to fuse the image features to obtain fused image features of the target device image;
[0024] an identification unit, configured to identify at least one device region containing device parts in the target device image based on the fused image features;
[0025] A classification unit is used to classify the device area based on the fused image features, and obtain device part information corresponding to the type of the device area to obtain device area identification information of the target device image.
[0026] A sending unit is configured to send the device area identification information to the terminal so that the terminal performs device area detection on the target device image based on the device area identification information.
[0027] Optionally, in some embodiments, the image recognition device may further include a fourth display unit, which may be specifically used to display part attribute information of the device area when the number of detected device areas does not exceed the preset number threshold, and the part attribute information includes attribute information of at least one device part; when the device area is not detected, prompt information is displayed on the part recognition page, and the prompt information is used to prompt the re-acquisition of the target device image.
[0028] Optionally, in some embodiments, the third display unit can be specifically used to display the area image corresponding to the target device area and the identification information of the device parts in the image area in the image area; and display the parts list of the target device area in the list area.
[0029] Optionally, in some embodiments, the third display unit can be specifically used to copy the attribute information in response to a trigger operation of a copy control for the attribute information; display a spare parts retrieval page in response to a trigger operation of the spare parts retrieval control, the spare parts retrieval page including an attribute information input area; add the copied attribute information to the attribute information input area, and display the spare parts information of the equipment parts corresponding to the attribute information on the spare parts retrieval page.
[0030] Optionally, in some embodiments, the third display unit can be specifically used to display an identification record query page in response to a triggering operation of the identification record query control, wherein the identification record query page includes a historical identification record list, and the historical identification record list includes attribute information of at least one historically identified equipment part.
[0031] Optionally, in some embodiments, the first display unit can be specifically used to display a user operation page of the client, where the user operation page includes a part identification control; and in response to a trigger operation on the part identification control, the part identification page is displayed.
[0032] Optionally, in some embodiments, the detection unit can be specifically used to obtain a target device image and send the target device image to a server for identification; receive device area identification information for the target device image returned by the server; and perform device area detection on the target device image based on the device area identification information.
[0033] Optionally, in some embodiments, the fusion unit can be specifically used to perform convolution processing on the image features using a trained recognition model to obtain convolved image features; adjust the size of the convolved image features according to the size of the image features; and fuse the adjusted image features to obtain fused image features of the target device image.
[0034] Optionally, in some embodiments, the image recognition device may further include a training unit, which may be specifically used to obtain positive samples of device images and non-device image samples, and use the positive samples of device images to train a preset recognition model to obtain an initial trained recognition model; use the initial trained recognition model to recognize the non-device image samples, and based on the recognition results, screen out incorrectly recognized non-device image samples from the non-device image samples to obtain negative samples of device images; use the positive samples and negative samples of device images to correct the initial trained recognition model to obtain the trained recognition model.
[0035] Optionally, in some embodiments, the recognition unit can be specifically used to extract the device area features of the device area from the fused image features; determine the device area information in the target device image based on the device area features; and identify the location information of at least one device area in the target device image based on the device area information to obtain at least one device area containing device parts in the target device image.
[0036] Optionally, in some embodiments, the classification unit can be specifically used to extract image features corresponding to the device area of a preset size from the fused image features to obtain target image features; obtain pooling weights of the target image features according to the size of the device area, and weight the target image features based on the pooling weights; pool the weighted image features, and determine the type of the corresponding device area based on the pooled image features.
[0037] In addition, an embodiment of the present invention further provides an electronic device, including a processor and a memory, wherein the memory stores an application program, and the processor is configured to run the application program in the memory to implement the image recognition method provided by the embodiment of the present invention.
[0038] In addition, an embodiment of the present invention further provides a computer-readable storage medium, which stores a plurality of instructions suitable for loading by a processor to execute the steps in any one of the image recognition methods provided by the embodiments of the present invention.
[0039] After displaying the part identification page of the client, the embodiment of the present invention responds to the trigger operation of the image acquisition control, performs device area detection on the acquired target device image, and when the number of detected device areas exceeds a preset number threshold, displays an area selection page, and responds to the selection operation of the device area, displays a part identification result page, which includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area; because the scheme acquires the target device image through the user's trigger operation of the image acquisition control, and then performs device area detection on the target device image, when the number of detected device areas exceeds a preset number threshold, the area selection page is displayed, so that the user can interactively select the target device area, and the part identification results in the target area are displayed, thereby realizing fast and accurate identification of the device parts of the target device, thereby improving the accuracy of image recognition. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0041] Figure 1 is a schematic diagram of an image recognition system provided by an embodiment of the present invention;
[0042] Figure 2 Schematic diagram of a scenario of an image recognition method provided by an embodiment of the present invention;
[0043] Figure 3 1 is a flow chart of an image recognition method provided by an embodiment of the present invention;
[0044] Figure 4 is a schematic diagram of a user operation page provided by an embodiment of the present invention;
[0045] Figure 5 1 is a schematic diagram of a parts identification page provided by an embodiment of the present invention;
[0046] Figure 6 is a schematic diagram of a device area in a target device image provided by an embodiment of the present invention;
[0047] Figure 7 is a schematic diagram of a region selection page provided by an embodiment of the present invention;
[0048] Figure 8 1 is a schematic diagram of a parts display page provided by an embodiment of the present invention;
[0049] Figure 9 1 is a schematic diagram of a part recognition result page provided by an embodiment of the present invention;
[0050] Figure 10 This is an identification record query page provided by an embodiment of the present invention;
[0051] Figure 11 is another flowchart of the image recognition process provided by an embodiment of the present invention;
[0052] Figure 12 is a schematic diagram of local area detection of a target recognition image provided by an embodiment of the present invention;
[0053] Figure 13 is another flowchart of the image recognition process provided by an embodiment of the present invention;
[0054] Figure 14 1 is a schematic diagram of the recognition process of image recognition provided by an embodiment of the present invention;
[0055] Figure 15 is a schematic diagram of visually displaying image recognition results provided by an embodiment of the present invention;
[0056] Figure 16 is a structural diagram of a first image recognition device provided by an embodiment of the present invention;
[0057] Figure 17 is another structural schematic diagram of the first image recognition device provided by an embodiment of the present invention;
[0058] Figure 18 is a structural diagram of a second image recognition device provided by an embodiment of the present invention;
[0059] Figure 19 is another structural diagram of a second image recognition device provided by an embodiment of the present invention;
[0060] Figure 20 It is a structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0062] The embodiments of the present invention provide an image recognition method, apparatus, electronic device, and computer-readable storage medium. The image recognition apparatus can be integrated into an electronic device, which can be a server, a terminal, or other device.
[0063] Specifically, the embodiment of the present application provides an image recognition device suitable for a first electronic device (for the purpose of distinction, it can be referred to as a first image recognition device), and an image recognition device suitable for a second electronic device (for the purpose of distinction, it can be referred to as a second image recognition device). Among them, the first electronic device can be a terminal or other device, and the terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The second electronic device can be a network side device such as a server, and the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal and the server can be directly or indirectly connected via wired or wireless communication, and this application does not limit this.
[0064] In the embodiment of the present application, the first electronic device is taken as a terminal and the second electronic device is taken as a server. For example, see Figure 1 The image recognition system provided by the embodiment of the present invention includes a terminal 10 and a server 20, wherein the terminal 10 and the server 20 are connected via a network, for example, via a wired or wireless network.
[0065] The terminal 10 may obtain the device region identification information of the target device image from the server 20, and perform device region detection on the target device image based on the device region identification information, such as Figure 2 As shown, the specific details can be as follows:
[0066] A parts identification page of the client is displayed, which includes an image acquisition control. In response to a trigger operation on the image acquisition control, device area detection is performed on the acquired target device image. When the number of detected device areas exceeds a preset threshold, an area selection page is displayed. The area selection page includes multiple device areas to be selected. In response to a selection operation on the device area, a parts identification result page is displayed. The part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area.
[0067] Among them, the response is used to indicate the conditions or states on which the executed operations depend. When the dependent conditions or states are met, one or more operations executed can be real-time or have a set delay; unless otherwise specified, there is no restriction on the order of execution of the multiple operations executed.
[0068] The server 20 is configured to receive the target device image sent by the terminal, identify the target device image, and send device area identification information of the target device image to the terminal 10. Specifically, the server 20 may be configured as follows:
[0069] A target device image sent by a receiving terminal is subjected to local feature extraction on the target device image to obtain image features of multiple sizes. The image features are then fused to obtain fused image features of the target device image. Based on the fused image features, at least one device area containing device parts is identified in the target device image. The device area is then classified based on the fused image features, and device part information corresponding to the type of the device area is obtained to obtain device area identification information of the target device image. The device area identification information is then sent to the terminal so that the terminal can perform device area detection on the target device image based on the device area identification information.
[0070] Among them, the image recognition method provided in the embodiment of the present application involves computer vision technology in the field of artificial intelligence, that is, in the embodiment of the present application, the computer vision technology of artificial intelligence can be used to extract features of the target device image, and the device area identification information of the target device image can be determined based on the extracted features, and the image can be recognized based on the device area identification information.
[0071] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. AI software technologies primarily include computer vision and machine learning / deep learning.
[0072] Computer vision (CV) is the science of making machines "see." Specifically, it refers to the use of computers to identify and measure objects, replacing the human eye. Image processing is then performed to make the image more suitable for human observation or for transmission to instrumentation. As a scientific discipline, computer vision studies related theories and technologies, aiming to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technology typically includes techniques such as image processing and image recognition, as well as common biometric recognition technologies such as facial recognition and human posture recognition.
[0073] The trained recognition model can be deployed on a cloud platform, and device region identification information can also be stored on the cloud platform. A cloud platform, also known as a cloud computing platform, refers to services based on hardware and software resources that provide computing, networking, and storage capabilities. Cloud computing is a computing model that distributes computing tasks across a resource pool consisting of a large number of computers, enabling various application systems to access computing power, storage space, and information services as needed. The network that provides resources is called the "cloud." To users, resources in the "cloud" appear infinitely scalable and can be accessed at any time, used on demand, expanded at any time, and paid for on a pay-per-use basis.
[0074] As a provider of cloud computing infrastructure, a cloud computing resource pool (referred to as a cloud platform, generally referred to as an IaaS (Infrastructure as a Service) platform) is established. Various types of virtual resources are deployed in the resource pool for external customers to choose and use. The cloud computing resource pool mainly includes: computing devices (virtualized machines, including operating systems), storage devices, and network devices.
[0075] Based on logical functional divisions, the PaaS (Platform as a Service) layer can be deployed on top of the IaaS (Infrastructure as a Service) layer, and the SaaS (Software as a Service) layer can be deployed on top of the PaaS layer. SaaS can also be deployed directly on top of IaaS. PaaS is a platform for software execution, such as databases and web containers. SaaS is a variety of business software, such as web portals and text messaging apps. Generally speaking, SaaS and PaaS are layers above IaaS.
[0076] It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.
[0077] This embodiment will be described from the perspective of an image recognition device (i.e., a first image recognition device), which can specifically be integrated with a terminal and other devices; wherein, the terminal can include a tablet computer, a laptop computer, a personal computer (PC), a wearable device, a virtual reality device or other smart devices that can perform image recognition.
[0078] An image recognition method, comprising:
[0079] A part identification page of the client is displayed, which includes an image acquisition control. In response to a trigger operation on the image acquisition control, device area detection is performed on the acquired target device image. When the number of detected device areas exceeds a preset threshold, an area selection page is displayed. The area selection page includes multiple device areas to be selected. In response to a selection operation on the device area, a part identification result page is displayed. The part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area.
[0080] like Figure 3 As shown, the specific process of the image recognition method is as follows:
[0081] 101. Display the parts identification page.
[0082] The part identification page can be understood as a page for acquiring device images for part identification. The part identification page includes an image acquisition control, which can be used to acquire device images. Image acquisition controls can be of various types, such as image acquisition controls or image filtering controls. These controls can take various forms, such as input boxes, icons, and buttons.
[0083] There are many ways to display the parts identification page, which can be as follows:
[0084] For example, a user operation page of the client may be displayed, where the user operation page includes a parts identification control, and in response to a trigger operation on the parts identification control, the parts identification page is displayed.
[0085] There are many ways to display the user operation page of the client. For example, the user can display the user operation page of the client by triggering the workbench control of the client on the terminal. The client can also be in various forms, for example, it can be an instant messaging client, or a business client, etc. The business type of the business client can be in various forms. For example, taking the business of spare parts retrieval as an example, the user operation page can be as follows: Figure 4shown.
[0086] After the user operation page is displayed, the user can trigger the parts identification control to display the parts identification page. There are also many ways to display the parts identification page. For example, the user's identity can be obtained, and based on the identity, the user's usage time of the parts identification service and the commonly used spare parts information can be determined. Then, based on the usage time and commonly used spare parts, a parts identification page is generated and displayed on the client. The displayed parts identification page can be as follows: Figure 5 As shown, the image acquisition control can include image collection controls and image filtering controls.
[0087] 102. In response to a triggering operation on the image acquisition control, perform device area detection on the acquired target device image.
[0088] Among them, device area detection can be understood as detecting whether the target device image contains a device area. The so-called device area can be understood as a part of the device or an area containing one or more parts. Since the device usually has many large and small parts, in order to better describe the device, the device is usually divided into multiple parts (areas) according to the positional relationship of the parts, etc., and these parts (areas) are used as the device area of the device. Specifically, Figure 6 As shown, generally, a device may include multiple device areas, but a captured device image may include one or more device areas, or may not include any device area.
[0089] There are multiple ways to perform device area detection on the target device image, which are as follows:
[0090] For example, a target device image is obtained and sent to a server for identification, device region identification information for the target device image returned by the server is received, and device region detection is performed on the target device image based on the device region identification information.
[0091] There are multiple ways to acquire the target device image. For example, when the image acquisition control is an image capture control, in response to a trigger operation on the image capture control, the terminal's image capture device is called to capture the target device image. The image capture device can be of various types, such as a front camera, a rear camera, or an independent camera. When the image acquisition control is an image filtering control, in response to a trigger operation on the image filtering control, an image selection page is displayed. The image selection page includes at least one candidate device image. In response to a selection operation on a candidate device image, the candidate device image selected by the selection operation is used as the target device image.
[0092] After acquiring the target device image, the target device image can be subjected to device area detection. There are many ways to detect the device area. For example, the device area identification information can be queried to determine whether the device area exists. Alternatively, the device area identification information can be queried to determine whether part attribute information exists. When part attribute information exists, the location information of the spare parts is extracted from the part attribute information. Based on the location information, the number and type of device areas existing in the target device image are determined, thereby completing the device area detection.
[0093] 103. When the number of detected device areas exceeds a preset threshold, a region selection page is displayed.
[0094] The area selection page includes multiple device areas to be selected.
[0095] There are many ways to display the area selection page, which are as follows:
[0096] For example, when the number of detected device areas exceeds a preset threshold, the location information of the device area is extracted from the device area identification information, and based on the location information, the device area is marked in the target device image, thereby generating a region selection page. Then, the region selection page is displayed. The region selection page can be as follows: Figure 7 shown.
[0097] There are many ways to mark the device area in the target device image. For example, the device area can be identified in the target device image based on the position information, and then, based on the size information of the device area, an identification frame of the device area can be added to the target device image to mark the device area in the target device image. Alternatively, an identification frame can be generated based on the position information and added to the target device image to mark the corresponding device area.
[0098] The preset number threshold can be set according to actual application and can be any value, for example, 1, 2 or any number.
[0099] Optionally, when the detected device area does not exceed the preset quantity threshold, the part attribute information of the device area is displayed. The part attribute information includes the attribute information of at least one device part. The attribute information of the device part may include information such as name, number, specification, and size. For example, taking the preset quantity threshold as 1, that is, when the number of detected device areas is 1, the part attribute information contained in the device area can be extracted from the device area identification information, and the part attribute information can be directly displayed on the terminal, or, based on the part attribute information, a part display page can be generated, and the part attribute information of at least one part contained in the device area can be displayed on the part display page. Specifically, Figure 8 shown.
[0100] Optionally, when the device area is not detected, a prompt message is displayed on the part identification page, prompting the user to reacquire the target device image. When the user sees this prompt message, the image acquisition control on the part identification page can be triggered. In response to the triggering operation on the image acquisition control, the target device image is reacquired and the device area detection is performed on the target device image.
[0101] It should be noted that the failure to detect the device area simply means that the acquired target device image does not contain the device area, or does not completely contain a device area. In this case, the user will be prompted to reacquire the target device image.
[0102] 104. In response to the selection operation on the equipment area, a parts identification result page is displayed.
[0103] The part recognition result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area. The part recognition result page also includes a list area and an image area, which can be specifically as follows: Figure 9 shown.
[0104] There are many ways to display the part recognition result page, which can be as follows:
[0105] For example, the image area displays the area image corresponding to the target device area and identification information of the device parts in the image area, and the list area displays the parts list of the target device area.
[0106] There are various ways to display the region image and the identification information of the device parts. For example, the image within the device region can be cropped from the target device image to obtain a region image corresponding to the target device region, and then the region image can be displayed in the image region. Regarding the identification information of the device parts, the part identification information corresponding to the device region can be extracted from the device region identification information. Based on the part identification information, the part location information of the device parts contained in the target device region can be identified in the region image. Based on the part location information, an identification frame for the device parts can be generated, and the identification frame for the device parts can be added to the region image to obtain the identification information of the device parts.
[0107] There are many ways to display the parts list of the target device area in the list area. For example, the attribute information of the device parts contained in the target device area can be extracted from the device area identification information, and the attribute information of the device parts can be added to the preset parts list template to obtain the parts list corresponding to the target device area, and the parts list can be displayed in the list area.
[0108] Optionally, the part identification results page also includes a spare parts search control and a copy control for each device part's attribute information. Therefore, after displaying the part identification results page, the identified device part's attribute information can be used to retrieve spare parts information for the device part. For example, in response to a triggering operation on the attribute information copy control, the attribute information is copied. In response to a triggering operation on the spare parts search control, a spare parts search page is displayed. The spare parts search page includes an attribute information input area, and the copied attribute information is added to the attribute information input area. The spare parts search page then displays the spare parts information for the device part corresponding to the attribute information. Spare parts information here can be understood as information about the spare parts of the device part on the target device. Spare parts can be understood as replacement parts for device parts. When a device part has a problem, the parts used to replace or repair the device part can be called spare parts. Spare parts information can include various information, such as the spare part's storage address, model, size, specifications, and installation / repair / maintenance precautions.
[0109] Optionally, the part identification result page also includes an identification record query control for querying historical identification records. Therefore, after displaying the part identification page, the identification record query page can also be displayed. For example, in response to the triggering operation of the identification record query control, the identification record query page is displayed. Specifically, Figure 10 As shown, the identification record query page includes a historical identification record list, which includes attribute information of at least one historically identified device part. The identification record query page can also include a query condition setting control. The user sets the query condition by triggering the query condition setting control. The query condition can include multiple conditions, such as query time conditions and device area query conditions, etc.
[0110] From the above, it can be seen that after the embodiment of the present application displays the part identification page of the client, in response to the trigger operation of the image acquisition control, the device area detection is performed on the acquired target device image. When the number of detected device areas exceeds the preset number threshold, the area selection page is displayed. In response to the selection operation of the device area, the part identification result page is displayed. The part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area. Since the scheme acquires the target device image through the user's trigger operation of the image acquisition control, and then performs device area detection on the target device image, when the number of detected device areas exceeds the preset number threshold, the area selection page is displayed, so that the user can interactively select the target device area and display the part identification results in the target area, thereby realizing fast and accurate identification of the parts of the target device, thereby improving the accuracy of image recognition.
[0111] This embodiment will be described from the perspective of a second image recognition device, which can be specifically integrated into an electronic device, which can be a server or other device; wherein the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), as well as big data and artificial intelligence platforms.
[0112] An image recognition method, comprising:
[0113] A target device image sent by a receiving terminal is extracted locally on the target device image to obtain image features of multiple sizes, the image features are fused to obtain fused image features of the target device image, at least one device area containing device parts is identified in the target device image based on the fused image features, the device area is classified based on the fused image features, and device part information corresponding to the type of the device area is obtained to obtain device area identification information of the target device image, and the device area identification information is sent to the terminal so that the terminal can perform device area detection on the target device image based on the device area identification information.
[0114] like Figure 11 As shown, the specific process of the image recognition method is as follows:
[0115] 201. Receive a target device image sent by a terminal, and perform local feature extraction on the target device image to obtain image features of multiple sizes.
[0116] For example, a target device image sent by a receiving terminal is extracted, and a trained recognition model is used to extract local features of the target device image to obtain image features of multiple sizes. Specifically, the following can be performed:
[0117] (1) Receive the target device image sent by the terminal.
[0118] For example, the target device image sent by the terminal can be directly received, or, when the number of target device images is large or the memory is large, an image recognition request sent by the terminal can be received, which carries the storage address of the target device image, and the target device image is obtained according to the storage address.
[0119] (2) The trained recognition model is used to extract local features of the target device image to obtain image features of multiple sizes.
[0120] For example, a trained recognition model can be used to connect high-resolution and low-resolution networks in parallel to extract local features of the target device image, thereby obtaining high-resolution image features of multiple sizes.
[0121] Among them, the network structure of the post-training recognition model can be diverse, for example, a convolutional neural network (HRNet) can be used to maintain high-resolution spatial structure. HRNet connects high-resolution and low-resolution networks in parallel, so it can maintain high resolution and the predicted heatmap is more spatially accurate.
[0122] The image features of multiple sizes may be multi-level features with strong semantic information, and the sizes of the multi-level features increase from top to bottom.
[0123] 202. Fusing the image features to obtain fused image features of the target device image.
[0124] For example, the trained recognition model can be used to perform convolution processing on the image features to obtain convolved image features. According to the size of the image features, the size of the convolved image features is adjusted, and the adjusted image features are fused to obtain fused image features of the target device image.
[0125] Among them, there are many ways to perform convolution processing on image features. For example, taking multi-level features including {C2, C3, C4, C5} as an example, first perform convolution coding on C5 to generate the top feature P5, then perform convolution coding on the remaining {C2, C3, C4} to generate {P2, P3, P4}, and use {P2, P3, P4, P5} as the convolved image features.
[0126] After obtaining the convolution image features, the size of the convolution image features can be adjusted in many ways, for example, P i+1 (i=4,3,2) upsampled to C i The same size, that is, upsample P4 to the size of C3, upsample P3 to the size of C2, and upsample P5 to the size of C4. The goal of upsampling is to ensure that P i+1 and C i Maintain the same spatial resolution.
[0127] After adjusting the size of the convolved image features, the adjusted image features can be fused. There are many ways to fusion. For example, the fused image features can be fused by pixel-level addition, as shown in formula (1):
[0128] P i =Conv(C i )+UP(P i+1) (i=4,3,2) (1)
[0129] Among them, P i is the fused image feature of the local area, Conv is a 1x1 convolution operation, the purpose of which is to control C i The number of channels, UP is the upsampling operation, its purpose is to ensure P i+1 and C i Keeping the same spatial resolution, C i is the image feature, P i+1 Adjusted image features.
[0130] The trained recognition model can be set according to the needs of actual applications. In addition, it should be noted that the trained recognition model can be pre-set by maintenance personnel or trained by the image recognition device itself. That is, before the step of "using the trained recognition model to perform convolution processing on the image features to obtain convolved image features", the image recognition method may further include:
[0131] Positive device image samples and non-device image samples are obtained, and a preset recognition model is trained using the device image samples to obtain an initial trained recognition model. The initial trained recognition model is used to recognize non-device image samples, and based on the recognition results, incorrectly recognized non-device image samples are screened out from the non-device image samples to obtain negative device image samples. The positive device image samples and the negative device image samples are used to correct the initial trained recognition model to obtain a trained recognition model. Specifically, the following can be done:
[0132] (1) Obtain positive samples of device images and non-device image samples, and use the positive samples of device images to train a preset recognition model to obtain an initial trained recognition model.
[0133] Among them, the positive sample of the equipment image can be an image containing the equipment in the workshop to be detected Sample The image sample will be Y i P As a marking device area type and location annotation, etc. Non-device image samples are image samples of equipment not in the workshop to be inspected, image samples of non-equipment in the workshop to be inspected, or image samples unrelated to equipment.
[0134] There are many ways to obtain positive device image samples and non-device image samples. For example, positive device image samples and non-device image samples can be received from users, or they can be obtained from an image database. Alternatively, original images can be obtained from the Internet, and images of equipment containing the workshop to be inspected can be identified in the original images. At least one equipment area and the type corresponding to the equipment area are marked in the equipment image, thereby obtaining positive device image samples. Original images in which equipment containing the workshop to be inspected is not identified are used as non-device image samples.
[0135] After obtaining the positive samples of the device image, the preset recognition model can be trained using the positive samples of the device image. The training process can be carried out in various ways. For example, the preset recognition model can be used to predict the device area position and category of the positive sample of the device image, and the predicted device area position and category can be compared with the position and category of the marked device area to obtain the loss information corresponding to the positive sample of the device image. Based on the loss information, the network parameters of the preset recognition model can be updated to obtain the recognition model after initial training.
[0136] (2) The initial trained recognition model is used to identify non-device image samples, and based on the recognition results, the non-device images with incorrect recognition are screened out from the non-device image samples to obtain negative samples of device images.
[0137] For example, the recognition model after initial training is used to recognize non-device image samples, and the recognition results are compared with the annotation information in the non-device image samples. When the recognition results are different from the annotation information, it can be determined that the non-device image recognition is incorrect. When the recognition results are the same as the annotation information, it can be determined that the non-device image recognition is correct. Non-device images with incorrect recognition are screened out from the non-device images, and the screened non-device images are marked as other categories. Thus, we get the negative sample of the device image
[0138] (3) The initial trained recognition model is corrected using positive and negative device image samples to obtain a trained recognition model.
[0139] The process of correcting the initial training recognition model can also be considered a training process. The training image samples include positive device image samples and negative device image samples. There are many ways to correct the model, which can be as follows:
[0140] For example, the positive sample of the device image and device image negative samples The data is input into the initial trained recognition model so that the initial trained recognition model predicts the location and category of the device area in the positive sample and the negative sample of the device image, and the predicted device area and category are compared with the annotation information to obtain the loss information of the positive sample and the negative sample of the device image. Based on the loss information, the network parameters of the initial trained recognition model are updated to obtain the trained recognition model.
[0141] Among them, the trained recognition model corrected by the positive and negative device image samples not only has a good recognition ability for specific areas of the device, but also can avoid some unnecessary wrong predictions, such as other categories.
[0142] 203. Identify at least one device region containing device parts in the target device image based on the fused image features.
[0143] For example, the device area features of the device area can be extracted from the fused image features, and the device area information in the target device image can be determined based on the device area features. Based on the device area information, the location information of at least one device area in the target device image can be identified to obtain at least one device area containing device parts in the target device image.
[0144] The device region information may be information such as the number and size of prediction boxes indicating the candidate device region. There are many ways to determine the device region information in the target device image based on the device region features. For example, the device region features can be identified to obtain the predicted number n of candidate regions, and the final height and width of the candidate device region after feature alignment can be extracted from the device region features to be k, thus generating nk. 2 The number of prediction boxes and the size information of the candidate device area are used as the device area information.
[0145] After determining the device area information, the position information of at least one device area can be identified in the target device image, thereby obtaining at least one device area containing device parts in the target device image. There are many ways to identify it. For example, a dense local regression network can be used to predict the position of the prediction box in the target device image, thereby identifying the position information of at least one device area. The network structure of the dense local regression network can be various. For example, a fully convolutional network can be used to generate multiple position-sensitive box offsets, thereby predicting the position information of at least one device area.
[0146] 204. Based on the fused image features, classify the device area and obtain device part information corresponding to the type of the device area to obtain device area identification information of the target device image.
[0147] For example, there are many ways to classify device areas. For example, image features corresponding to a device area of a preset size can be extracted from the fused image features to obtain target image features. According to the size of the device area, the pooling weight of the target image features is obtained, and based on the pooling weight, the target image features are weighted. The weighted image features are pooled, and the type of the corresponding device area is determined based on the pooled image features.
[0148] Among them, the preset size can be set according to the actual application. Taking the size of the device area as k*k as an example, the preset size can be k / 2*K / 2, thereby realizing lightweight offset prediction. Compared with the standard offset prediction in the existing deformable RoI-Pooling, this solution only requires about one-fourth of the parameters.
[0149] After extracting the target image features, the pooling weights of the target image features can be obtained. The pooling weights here can be understood as the pooling weights of the four sampling points in the device area. This solution adopts weighted pooling to adaptively allocate different pooling weights.
[0150] Based on the obtained pooling weights, the features sampled at different sampling points in the target image features are weighted to obtain weighted image features, which are then pooled. Based on the pooled image features, the type of the corresponding device area is determined. This determination can be done in various ways, such as connecting the pooled image features corresponding to the device area to a fully connected layer to obtain the predicted probability of the candidate type corresponding to the device area. Based on the predicted probability, the type corresponding to the device area is selected from the candidate types.
[0151] After determining the type corresponding to the equipment area, the equipment parts information corresponding to the type can be obtained, thereby obtaining the equipment area identification information. For example, the equipment parts information corresponding to the type of the equipment area is filtered out from the preset equipment parts information set, and the equipment parts information is added to the equipment parts template corresponding to the type of the equipment area to obtain the equipment parts information corresponding to the equipment area. The equipment parts information of each identified equipment area is integrated to obtain the equipment area identification information.
[0152] 205. Send the device region identification information to the terminal, so that the terminal performs device region detection on the target device image based on the device region identification information.
[0153] For example, the device area identification information can be sent directly to the terminal, or, when the memory of the device area information is large or the amount is large, the device area identification information can be stored, a storage address is obtained, and the storage address is sent to the terminal, so that the terminal obtains the device area identification information according to the storage address.
[0154] After receiving the device region identification information, the terminal can perform device region detection on the target device image based on the device region identification information. The detection process is described above and will not be described in detail here.
[0155] It should be noted that the recognition of the target device image in this solution is to detect the local area of the target device image to realize the recognition of the location and type of the device area. The specific process can be as follows: Figure 12 As shown, given a target device image, a convolutional neural network is first used to extract features. Then, a feature pyramid network is used to encode information of different scales and levels. Finally, a discriminative local feature alignment network and a dense regression network are used to realize the positioning and category recognition of the device area (target). The entire recognition process can be regarded as the recognition of the target device image through the trained recognition model. Regarding the detection accuracy and recall rate of the device area, on a test set of 6974 machine tool locations, under the conditions of accurate identification of the location and IOU>=0.5, the v2 algorithm in this solution predicts an overall accuracy of 97.26% and a recall rate of 95.81%. Compared with the basic version (v1) algorithm, using the same test set and statistical method, the overall accuracy of this solution is improved by 2.85% and the recall rate is improved by 9.43%, as shown in Table 1:
[0156] Table 1
[0157]
[0158]
[0159] False detection rate for non-device images:
[0160] (a) False detection test A: 1W external test set (non-device images), the basic version (v1) falsely detected more than 4000 images (false detection rate 40%+), and the present invention (v2) falsely detected 0 images (false detection rate 0%);
[0161] (b) False detection test B: 1 million images of other projects (40w for UGC, 30w for live broadcast, and 30w for on-demand broadcast, all of which are non-machine tool images), the present invention (v2) falsely detected 184 images (false detection rate 0.02%);
[0162] (c) False detection test C: 596 photos of non-equipment parts of the factory environment were used, and 3 photos were falsely detected (false detection rate 0.33%).
[0163] We also tested the impact of compressed images on the recognition performance of the trained recognition model. We compressed the original test images into JPEGs at compression rates of 10%, 30%, 50%, and 75% to verify the performance of the detection algorithm. Table 2 shows consistent performance, indicating that the detection algorithm is not very sensitive to compression rates and is very robust. This suggests that the proposed solution is applicable to mobile phones with a variety of different sensors.
[0164] Table 2 - Compression ratio (10%)
[0165]
[0166]
[0167] Compression rate (30%)
[0168] Region Name gt Pred OK Accuracy Recall Area 1 2032 1995 1944 97.44% 95.67% Area 2 952 900 897 99.67% 94.22% Area 3 1063 1057 1012 95.74% 95.20% Area 4 1887 1867 1840 98.55% 97.51% Area 5 1173 1189 1144 96.22% 97.53% Area 6 524 484 435 89.88% 83.02% Area 7 1016 961 918 95.53% 90.35% Area 8 1869 1832 1703 92.96% 91.12% overall 10516 10285 9893 96.19% 94.08%
[0169] Compression rate (50%)
[0170]
[0171]
[0172] Compression rate (75%)
[0173] Region Name gt Pred OK Accuracy Recall Area 1 2032 1995 1944 97.44% 95.67% Area 2 952 900 897 99.67% 94.22% Area 3 1063 1057 1012 95.74% 95.20% Area 4 1887 1867 1840 98.55% 97.51% Area 5 1173 1189 1144 96.22% 97.53% Area 6 524 484 435 89.88% 83.02% Area 7 1016 961 918 95.53% 90.35% Area 8 1869 1832 1703 92.96% 91.12% overall 10516 10285 9893 96.19% 94.08%
[0174] As can be seen from the above, this embodiment receives the target device image sent by the terminal, extracts local features of the target device image, obtains image features of multiple sizes, and then fuses the image features to obtain fused image features of the target device image. Then, based on the fused image features, at least one device area containing device parts is identified in the target device image. Based on the fused image features, the device area is classified, and device part information corresponding to the type of the device area is obtained to obtain device area identification information of the target device image. The device area identification information is sent to the terminal so that the terminal performs device area detection on the target device image based on the device area identification information. Since this scheme extracts local features of the target device image, then fuses the extracted image features, identifies at least one device area based on the fused image features, and determines the type of the device area, thereby obtaining device area identification information of the target device image, and sending the device area identification information to the terminal, the user can interactively select the target device area through the terminal, thereby quickly and accurately identifying the device parts of the target device, thereby improving the accuracy of image recognition.
[0175] The method described in the above embodiment will be further described in detail below with examples.
[0176] In this embodiment, the first image recognition device is used as a terminal, the second image recognition device is used as a server, and the target device image is used as a target machine tool image.
[0177] (1) The server trains the preset recognition model to obtain a trained recognition model.
[0178] (1) The server obtains positive samples of machine tool images and non-equipment image samples, and uses the positive samples of equipment images to train the preset recognition model to obtain the initial trained recognition model.
[0179] For example, the server can receive positive machine tool image samples and non-machine tool image samples uploaded by users, or can obtain positive machine tool image samples and non-machine tool image samples from an image database, or can obtain original images from the Internet, identify the machine tool image containing the workshop to be inspected in the original image, and annotate at least one equipment area and the type corresponding to the equipment area in the machine tool image, thereby obtaining positive machine tool image samples. Original images in which the machine tool containing the workshop to be inspected is not identified are used as non-machine tool image samples.
[0180] The server uses a preset recognition model to predict the device area location and category of the positive sample of the machine tool image, compares the predicted device area location and category with the marked device area location and category, obtains the loss information corresponding to the positive sample of the machine tool image, and updates the network parameters of the preset recognition model based on the loss information to obtain the recognition model after initial training.
[0181] (2) The server uses the initial trained recognition model to identify non-machine tool image samples, and based on the recognition results, screens out non-machine tool images with incorrect recognition from the non-machine tool image samples to obtain negative machine tool image samples.
[0182] For example, the server uses the recognition model after initial training to identify non-machine tool image samples, and compares the recognition results with the annotation information in the non-machine tool image samples. When the recognition results are different from the annotation information, it can be determined that the non-machine tool image recognition is incorrect. When the recognition results are the same as the annotation information, it can be determined that the non-machine tool image recognition is correct. Non-machine tool images with incorrect recognition are screened out from the non-machine tool images, and the screened non-machine tool images are marked as other categories. Thus, we get the negative sample of machine tool image
[0183] (3) The server uses positive samples and negative samples of machine tool images to correct the initial trained recognition model to obtain a trained recognition model.
[0184] For example, the server sends the positive sample of the machine tool image and machine tool image negative samples The data is input into the initial trained recognition model so that the initial trained recognition model predicts the location and category of the equipment area in the positive and negative machine tool image samples, and compares the predicted equipment area and category with the annotation information to obtain the loss information of the positive and negative machine tool image samples. Based on the loss information, the network parameters of the initial trained recognition model are updated to obtain the trained recognition model.
[0185] (2) The server uses the trained recognition model to recognize the target machine tool image sent by the terminal, and sends the recognized device area identification information to the terminal, so that the terminal can perform device area detection on the target machine tool image based on the device area identification information, etc.
[0186] like Figure 13 As shown, an image recognition method, the specific process is as follows:
[0187] 301. The terminal displays a parts identification page.
[0188] For example, the user can trigger the client's workbench control on the terminal to display the client's user operation page, respond to the trigger operation on the parts identification control, obtain the user's identity, and determine the user's usage time and commonly used spare parts information of the parts identification service based on the identity. Then, based on the usage time and commonly used spare parts, generate a parts identification page and display the parts identification page on the client.
[0189] 302. In response to a trigger operation on an image acquisition control, the terminal sends the acquired target machine tool image to a server.
[0190] For example, when the image acquisition control is an image collection control, the terminal, in response to a trigger operation on the image collection control, calls the terminal's image acquisition device to capture the target machine tool image. When the image acquisition control is an image filtering control, the terminal, in response to a trigger operation on the image filtering control, displays an image selection page including at least one candidate machine tool image. In response to a selection operation on a candidate machine tool image, the candidate machine tool image selected by the selection operation is used as the target machine tool image. The terminal then sends the target machine tool image to the server for recognition.
[0191] 303. The server receives the target machine tool image sent by the terminal, and performs local feature extraction on the target machine tool image to obtain image features of multiple sizes.
[0192] For example, the server can directly receive the target machine tool image sent by the terminal, or, when the number of target machine tool images is large or the memory is large, it can also receive an image recognition request sent by the terminal, which carries the storage address of the target machine tool image, and obtain the target machine tool image according to the storage address.
[0193] The server uses a high-resolution spatial structure-preserving convolutional neural network (HRNe) to extract local features of the target machine tool image, thereby obtaining high-resolution image features of multiple sizes.
[0194] 304. The server fuses the image features to obtain fused image features of the target machine tool image.
[0195] For example, taking the multi-level features including {C2, C3, C4, C5} as an example, the server first performs convolution coding on C5 to generate the top feature P5, and then performs convolution coding on the remaining {C2, C3, C4} to generate {P2, P3, P4}, and uses {P2, P3, P4, P5} as the convolution image features. i+1 (i=4,3,2) upsampled to C i The same size, that is, upsample P4 to the size of C3, upsample P3 to the size of C2, and upsample P5 to the size of C4. The fused image features are added and fused at the pixel level, as shown in formula (1).
[0196] 305. The server identifies at least one equipment area containing equipment parts in the target machine tool image based on the fused image features.
[0197] For example, the server can extract the device area features of the device area from the fused image features, identify the device area features, and thus obtain the predicted number n of candidate areas. The final height and width of the candidate device area after feature alignment extracted from the device area features are k, and then nk is generated. 2 The number of prediction boxes and the size information of the candidate equipment area are used as the equipment area information. A dense local regression network is used to predict the position of the prediction box in the target machine tool image, thereby identifying the position information of at least one equipment area.
[0198] 306. The server classifies the equipment area based on the fused image features and obtains equipment part information corresponding to the type of the equipment area to obtain equipment area identification information of the target machine tool image.
[0199] For example, the server can extract image features corresponding to a device region of a preset size from the fused image features to obtain target image features. Based on the size of the device region, the server can obtain pooling weights for the target image features, weight the target image features based on the pooling weights, and perform pooling on the weighted image features. The pooled image features corresponding to the device region are connected to a fully connected layer to obtain predicted probabilities for candidate types corresponding to the device region. Based on the predicted probabilities, the type corresponding to the device region is selected from the candidate types.
[0200] The server filters out the equipment parts information corresponding to the type of the equipment area from the preset equipment parts information set, adds the equipment parts information to the equipment parts template corresponding to the type of the equipment area, thereby obtaining the equipment parts information corresponding to the equipment area, and merges the equipment parts information of each identified equipment area to obtain the equipment area identification information.
[0201] 307. The server sends the device region identification information to the terminal.
[0202] For example, the server can directly send the device area identification information to the terminal, or, when the memory of the device area information is large or the amount is large, the device area identification information can also be stored, a storage address is obtained, and the storage address is sent to the terminal, so that the terminal obtains the device area identification information according to the storage address.
[0203] 308. The terminal performs device area detection on the target machine tool image based on the device area identification information.
[0204] For example, the server can query whether there is an equipment area in the equipment area identification information, or it can also query whether there is part attribute information in the equipment area identification information. When part attribute information exists, the location information of the spare parts is extracted from the part attribute information. Based on the location information, the number and type of equipment areas existing in the target machine tool image are determined, thereby completing the equipment area detection.
[0205] 309. When the number of detected device areas exceeds a preset threshold, the terminal displays an area selection page.
[0206] For example, when the number of detected device areas exceeds a preset number threshold, the terminal extracts the location information of the device area from the device area identification information, identifies the device area in the target machine tool image based on the location information, and then adds an identification frame of the device area to the target machine tool image based on the size information of the device area, thereby marking the device area in the target machine tool image. Alternatively, an identification frame may be generated based on the location information, and added to the target machine tool image, and the corresponding device area may be marked by the identification frame, thereby generating an area selection page, and then displaying the area selection page.
[0207] Optionally, when the detected equipment area does not exceed a preset quantity threshold, the terminal extracts the part attribute information contained in the equipment area from the equipment area identification information and directly displays the part attribute information on the terminal, or, based on the part attribute information, generates a part display page and displays the part attribute information of at least one part contained in the equipment area on the part display page.
[0208] Optionally, when the device area is not detected, the terminal displays a prompt on the part identification page, prompting the user to reacquire the target machine tool image. When the user sees this prompt, they can trigger the image acquisition control on the part identification page. In response to the triggering operation on the image acquisition control, the target machine tool image is reacquired and the device area detection is performed on the target machine tool image.
[0209] 310. In response to the selection operation on the device area, the terminal displays a part identification result page.
[0210] For example, the terminal crops the image within the equipment area within the target machine tool image to obtain a regional image corresponding to the target equipment area, and then displays this regional image within the image area. Regarding the identification information of equipment parts, the terminal extracts the part identification information corresponding to the equipment area from the equipment area identification information. Based on this part identification information, the terminal identifies the part location information of the equipment parts within the target equipment area within the regional image. Based on this part location information, an identification frame for the equipment part is generated and added to the regional image to obtain the identification information of the equipment part.
[0211] The terminal extracts the attribute information of the equipment parts contained in the target equipment area from the equipment area identification information, adds the attribute information of the equipment parts to the preset parts list template, thereby obtaining the parts list corresponding to the target equipment area, and displays the parts list in the list area.
[0212] Optionally, the parts identification result page also includes a spare parts retrieval control and a copy control for the attribute information of each device part. After displaying the parts identification result page, the terminal can also, in response to a trigger operation on the copy control for the attribute information, copy the attribute information, and in response to a trigger operation on the spare parts retrieval control, display a spare parts retrieval page, which includes an attribute information input area, adds the copied attribute information to the attribute information input area, and displays the spare parts information of the device part corresponding to the attribute information on the spare parts retrieval page.
[0213] Optionally, the part identification result page also includes an identification record query control. After displaying the part identification page, the terminal can display an identification record query page in response to the triggering operation of the identification record query control. The identification record query page includes a historical identification record list. The historical identification record list includes attribute information of at least one historically identified device part. The identification record query page can also include a query condition setting control. The user sets the query condition by triggering the query condition setting control.
[0214] Among them, the part recognition in the target machine tool image can be regarded as completed by the interaction between the terminal and the server. The entire image recognition process can be described as follows: Figure 14 As shown, the customer maintenance staff (user) opens the parts identification or spare parts retrieval applet on the client, selects the photo identification function, and the user takes a photo to obtain the target machine tool image. After confirming that the equipment area detection is to be performed, the applet sends a request to the detection algorithm microservice. The algorithm (trained recognition model) detects all parts (equipment areas) in the target machine tool image and marks the category, and returns the detection results. The applet obtains the detection results, parses the position and code of the part (equipment area), and displays it visually to the user. If the part (equipment area) cannot be detected, the user is prompted to retake the photo. If a part is detected, the corresponding part in the material library is matched and the spare parts information is displayed. If multiple parts (equipment areas) are detected, the matching material library displays all detected parts (equipment areas), and the user is prompted to select one of the parts (equipment areas). The spare parts information is displayed after the user selects and confirms. The visual display of image recognition results (equipment area and spare parts information contained in the equipment area) can be as follows Figure 15 As shown, it can be seen that the use of this solution can accurately detect the specific parts of the machine tool (equipment area), especially in multiple departments, extreme shooting angles, exposure and other conditions.
[0215] From the above, it can be seen that after the embodiment of the present application displays the part identification page of the client, in response to the trigger operation of the image acquisition control, the device area detection is performed on the acquired target device image; when the number of detected device areas exceeds the preset number threshold, the area selection page is displayed, and in response to the selection operation of the device area, the part identification result page is displayed, and the part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area; because the scheme acquires the target device image through the user's trigger operation of the image acquisition control, and then performs device area detection on the target device image, when the number of detected device areas exceeds the preset number threshold, the area selection page is displayed, so that the user can interactively select the target device area and display the part identification results in the target area, thereby realizing fast and accurate identification of the parts of the target device, thereby improving the accuracy of image recognition.
[0216] In order to better implement the above method, an embodiment of the present invention further provides an image recognition device (i.e., a first image recognition device), which can be integrated into a terminal, which can include a smart phone, a tablet computer, a laptop computer and / or a personal computer, etc.
[0217] For example, Figure 16 As shown, the first image recognition device may include a first display unit 401, a detection unit 402, a second display unit 403 and a third display unit 404, as follows:
[0218] (1) a first display unit 401;
[0219] The first display unit 401 is used to display a part identification page of the client, where the part identification page includes an image acquisition control.
[0220] For example, the first display unit 401 may be specifically configured to display a user operation page of the client, where the user operation page includes a part identification control, and the part identification page is displayed in response to a triggering operation on the part identification control.
[0221] (2) Detection unit 402;
[0222] The detection unit 402 is configured to perform device area detection on the acquired target device image in response to a trigger operation on the image acquisition control.
[0223] For example, the detection unit 402 can be specifically used to obtain a target device image, send the target device image to a server for identification, receive device area identification information for the target device image returned by the server, and perform device area detection on the target device image based on the device area identification information.
[0224] (3) second display unit 403;
[0225] The second display unit 403 is configured to display an area selection page when the number of detected device areas exceeds a preset number threshold, where the area selection page includes a plurality of device areas to be selected.
[0226] For example, the second display unit 403 can be specifically used to extract the location information of the device area from the device area identification information when the number of detected device areas exceeds a preset number threshold, and based on the location information, mark the device area in the target device image, thereby generating an area selection page, and then displaying the area selection page.
[0227] (4) third display unit 404;
[0228] The third display unit 404 is used to display a part identification result page in response to a selection operation on a device area. The part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area.
[0229] For example, the third display unit 404 may be specifically configured to display an area image corresponding to the target device area and identification information of device parts in the image area in the image area, and to display a parts list of the target device area in the list area.
[0230] Optionally, the first image recognition device may further include a fourth display unit 405, such as Figure 17 As shown, the specific details can be as follows:
[0231] The fourth display unit 405 is used to display part attribute information in the equipment area or display prompt information on the part identification page.
[0232] For example, the fourth display unit 405 can be specifically used to display the part attribute information of the equipment area when the number of detected equipment areas does not exceed a preset number threshold. The part attribute information includes attribute information of at least one equipment part. When the equipment area is not detected, prompt information is displayed on the part identification page. The prompt information is used to prompt the re-acquisition of the target equipment image.
[0233] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.
[0234] As can be seen from the above, in this embodiment, after the first display unit 401 displays the client's part identification page, the detection unit 402 performs device area detection on the acquired target device image in response to the trigger operation on the image acquisition control. The second display unit 403 displays an area selection page when the number of detected device areas exceeds a preset number threshold. The third display unit 404 displays a part identification result page in response to the selection operation on the device area. The part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area. Since this scheme acquires the target device image through the user's trigger operation on the image acquisition control, and then performs device area detection on the target device image, when the number of detected device areas exceeds the preset number threshold, the area selection page is displayed, allowing the user to interactively select the target device area and display the part identification results in the target area, thereby achieving rapid and accurate identification of the device parts of the target device, thereby improving the accuracy of image recognition.
[0235] In order to better implement the above method, an embodiment of the present invention further provides an image recognition device (ie, a second image recognition device). The second image recognition device can be integrated into a server. The server can be a single server or a server cluster consisting of multiple servers.
[0236] For example, Figure 18 As shown, the second image recognition device may include a receiving unit 501, a fusion unit 502, a recognition unit 503, a classification unit 504 and a sending unit 505, as follows:
[0237] (1) receiving unit 501;
[0238] The receiving unit 501 is configured to receive a target device image sent by a terminal, and perform local feature extraction on the target device image to obtain image features of multiple sizes.
[0239] For example, receiving unit 501 can be specifically configured to directly receive a target device image sent by a terminal. Alternatively, when there are a large number of target device images or a large memory, it can also receive an image recognition request from the terminal, which carries the storage address of the target device image. The target device image is retrieved based on the storage address. A trained recognition model is used to connect high-resolution and low-resolution networks in parallel to extract local features from the target device image, thereby obtaining high-resolution image features at multiple sizes.
[0240] (2) fusion unit 502;
[0241] The fusion unit 502 is configured to fuse the image features to obtain fused image features of the target device image.
[0242] For example, the fusion unit 502 can be specifically used to use the trained recognition model to perform convolution processing on the image features to obtain the convolved image features, adjust the size of the convolved image features according to the size of the image features, and fuse the adjusted image features to obtain the fused image features of the target device image.
[0243] (3) Identification unit 503;
[0244] The recognition unit 503 is configured to recognize at least one device region containing device parts in the target device image based on the fused image features.
[0245] For example, the recognition unit 503 can be specifically used to extract the device area features of the device area from the fused image features, determine the device area information in the target device image based on the device area features, and identify the location information of at least one device area in the target device image based on the device area information to obtain at least one device area containing device parts in the target device image.
[0246] (4) classification unit 504;
[0247] The classification unit 504 is configured to classify the device area based on the fused image features, and obtain device part information corresponding to the type of the device area, thereby obtaining device area identification information of the target device image.
[0248] For example, classification unit 504 may be specifically configured to extract image features corresponding to a device region of a preset size from the fused image features to obtain target image features, obtain pooling weights for the target image features based on the size of the device region, weight the target image features based on the pooling weights, perform pooling on the weighted image features, and determine the type of the corresponding device region based on the pooled image features. Device part information corresponding to the type of the device region is obtained to obtain device region identification information for the target device image.
[0249] (5) sending unit 505;
[0250] The sending unit 505 is configured to send the device region identification information to the terminal so that the terminal can perform device region detection on the target device image based on the device region identification information.
[0251] For example, the sending unit 505 can be specifically used to directly send the device area identification information to the terminal, or, when the memory of the device area information is large or the amount is large, the device area identification information can also be stored to obtain a storage address, and the storage address is sent to the terminal, so that the terminal obtains the device area identification information according to the storage address.
[0252] Optionally, the second image recognition device may further include a training unit 506, such as Figure 19 As shown, the specific details can be as follows:
[0253] The training unit 506 is used to train the preset recognition model to obtain a trained recognition model.
[0254] For example, the training unit 506 can be specifically used to obtain positive samples of device images and non-device image samples, and use the device image samples to train a preset recognition model to obtain an initial trained recognition model, use the initial trained recognition model to recognize non-device image samples, and based on the recognition results, screen out non-device image samples with recognition errors from the non-device image samples to obtain negative samples of device images, and use the positive samples of device images and the negative samples of device images to correct the initial trained recognition model to obtain a trained recognition model.
[0255] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can be found in the previous method embodiments and will not be repeated here.
[0256] As can be seen from the above, in this embodiment, after the receiving unit 501 receives the target device image sent by the terminal and extracts local features of the target device image to obtain image features of multiple sizes, the fusion unit 502 fuses the image features to obtain fused image features of the target device image. Then, the recognition unit 503 identifies at least one device region containing device parts in the target device image based on the fused image features. The classification unit 504 classifies the device region based on the fused image features and obtains device part information corresponding to the type of the device region to obtain device region identification information of the target device image. The sending unit 505 sends the device region identification information to the terminal so that the terminal can perform device region detection on the target device image based on the device region identification information. Since this solution extracts local features from the target device image, then fuses the extracted image features, identifies at least one device region based on the fused image features, and determines the type of the device region, thereby obtaining device region identification information of the target device image, and sends the device region identification information to the terminal, the user can interactively select the target device region through the terminal, thereby quickly and accurately identifying the device parts of the target device. Therefore, the accuracy of image recognition can be improved.
[0257] An embodiment of the present invention further provides an electronic device, such as Figure 20 , which shows a schematic structural diagram of an electronic device involved in an embodiment of the present invention, specifically:
[0258] The electronic device may include one or more processing core processors 601, one or more computer-readable storage media memories 602, a power supply 603, an input unit 604 and other components. Those skilled in the art will understand that Figure 20 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.
[0259] The processor 601 is the control center of the electronic device, connecting the various parts of the entire electronic device using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 602 and calling data stored in the memory 602, it performs various functions of the electronic device and processes data, thereby performing overall detection and management of the electronic device. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly handles wireless communications. It is understood that the above-mentioned modem processor may not be integrated into the processor 601.
[0260] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.
[0261] The electronic device also includes a power supply 603 for supplying power to various components. Preferably, the power supply 603 can be logically connected to the processor 601 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 603 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0262] The electronic device may further include an input unit 604, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0263] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 601 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 602 according to the following instructions, and the processor 601 will run the application programs stored in the memory 602 to implement various functions as follows:
[0264] A part identification page of the client is displayed, which includes an image acquisition control. In response to a trigger operation on the image acquisition control, device area detection is performed on the acquired target device image. When the number of detected device areas exceeds a preset threshold, an area selection page is displayed. The area selection page includes multiple device areas to be selected. In response to a selection operation on the device area, a part identification result page is displayed. The part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area.
[0265] or
[0266] A target device image sent by a receiving terminal is extracted locally on the target device image to obtain image features of multiple sizes, the image features are fused to obtain fused image features of the target device image, at least one device area containing device parts is identified in the target device image based on the fused image features, the device area is classified based on the fused image features, and device part information corresponding to the type of the device area is obtained to obtain device area identification information of the target device image, and the device area identification information is sent to the terminal so that the terminal can perform device area detection on the target device image based on the device area identification information.
[0267] The specific implementation of the above operations can be found in the previous embodiments and will not be described in detail here.
[0268] From the above, it can be seen that after the embodiment of the present invention displays the part identification page of the client, in response to the trigger operation of the image acquisition control, the device area detection is performed on the acquired target device image, and when the number of detected device areas exceeds the preset number threshold, the area selection page is displayed, and in response to the selection operation of the device area, the part identification result page is displayed, and the part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area; because the scheme acquires the target device image through the user's trigger operation of the image acquisition control, and then performs device area detection on the target device image, when the number of detected device areas exceeds the preset number threshold, the area selection page is displayed, so that the user can interactively select the target device area and display the part identification results in the target area, thereby realizing fast and accurate identification of the parts of the target device, thereby improving the accuracy of image recognition.
[0269] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0270] To this end, an embodiment of the present invention provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the image recognition methods provided in the embodiments of the present invention. For example, the instructions can execute the following steps:
[0271] A part identification page of the client is displayed, which includes an image acquisition control. In response to a trigger operation on the image acquisition control, device area detection is performed on the acquired target device image. When the number of detected device areas exceeds a preset threshold, an area selection page is displayed. The area selection page includes multiple device areas to be selected. In response to a selection operation on the device area, a part identification result page is displayed. The part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area.
[0272] or
[0273] A target device image sent by a receiving terminal is extracted locally on the target device image to obtain image features of multiple sizes, the image features are fused to obtain fused image features of the target device image, at least one device area containing device parts is identified in the target device image based on the fused image features, the device area is classified based on the fused image features, and device part information corresponding to the type of the device area is obtained to obtain device area identification information of the target device image, and the device area identification information is sent to the terminal so that the terminal can perform device area detection on the target device image based on the device area identification information.
[0274] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0275] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0276] Since the instructions stored in the computer-readable storage medium can execute the steps in any image recognition method provided in the embodiments of the present invention, the beneficial effects that can be achieved by any image recognition method provided in the embodiments of the present invention can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0277] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform the methods provided in various optional implementations of the aforementioned image recognition or spare parts retrieval aspects.
[0278] The above is a detailed introduction to an image recognition method, device, electronic device and computer-readable storage medium provided in an embodiment of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for those skilled in the art, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. An image recognition method, characterized in that: include: Displaying a parts identification page of a client, wherein the parts identification page includes an image acquisition control; In response to a trigger operation on the image acquisition control, performing device area detection on the acquired target device image, wherein the device area detection is performed based on device area identification information returned by the server, the server identifying at least one device area containing device parts in the target device image, classifying the device area, determining a type corresponding to the device area, determining part information corresponding to the type of the device area from a preset device part information set based on the type of the device area, and obtaining the device area identification information based on the part information corresponding to the type of the device area; When the number of detected device areas exceeds a preset number threshold, a region selection page is displayed, wherein the region selection page includes multiple device areas to be selected; In response to a selection operation on the device area, a parts identification result page is displayed, the parts identification result page including a parts list of the target device area selected by the selection operation, the parts list including attribute information of at least one device part in the target device area.
2. The image recognition method according to claim 1, wherein: Also includes: When the number of detected equipment areas does not exceed the preset number threshold, displaying part attribute information of the equipment areas, the part attribute information including attribute information of at least one equipment part; When the device area is not detected, a prompt message is displayed on the part identification page, and the prompt message is used to prompt the user to re-acquire the target device image.
3. The image recognition method according to claim 1, wherein: The part recognition result page includes a list area and an image area. The part recognition result display page includes: Displaying, in the image area, an area image corresponding to the target device area and identification information of the device parts in the image area; A parts list of the target device area is displayed in the list area.
4. The image recognition method according to claim 3, wherein: The part identification result page also includes a spare parts search page and a copy control for the attribute information of each equipment part. After the part identification result page is displayed, it also includes: In response to a triggering operation of a copy control for the attribute information, copying the attribute information; In response to a trigger operation on the spare parts search control, displaying a spare parts search page, the spare parts search page including an attribute information input area; The copied attribute information is added to the attribute information input area, and the spare parts information of the equipment part corresponding to the attribute information is displayed on the spare parts search page.
5. The image recognition method according to claim 3, wherein: The part recognition result page also includes a recognition record query control, and after displaying the part recognition result page, it also includes: In response to the triggering operation of the identification record query control, an identification record query page is displayed, the identification record query page including a historical identification record list including attribute information of at least one historically identified device part.
6. The image recognition method according to any one of claims 1 to 5, characterized in that: The part identification page of the display client includes: Displaying a user operation page of the client, wherein the user operation page includes a part identification control; In response to a triggering operation on the part identification control, a part identification page is displayed.
7. The image recognition method according to any one of claims 1 to 5, characterized in that: The performing device area detection on the acquired target device image includes: Acquire a target device image and send the target device image to a server for identification; receiving device region identification information for the target device image returned by the server; Based on the device area identification information, device area detection is performed on the target device image.
8. An image recognition method, characterized in that: include: receiving a target device image sent by a terminal, and performing local feature extraction on the target device image to obtain image features of multiple sizes; fusing the image features to obtain fused image features of the target device image; identifying at least one device region containing device parts in the target device image based on the fused image features; Based on the fused image features, the device area is classified, and device part information corresponding to the type of the device area is obtained to obtain device area identification information of the target device image. The obtaining of the device part information corresponding to the type of the device area includes: determining, according to the type of the device area, the part information corresponding to the type of the device area from a preset device part information set; The device area identification information is sent to the terminal, so that the terminal performs device area detection on the target device image based on the device area identification information.
9. The image recognition method according to claim 8, characterized in that: The fusing the image features to obtain fused image features of the target device image includes: Performing convolution processing on the image features using the trained recognition model to obtain convolution image features; Adjusting the size of the convolved image feature according to the size of the image feature; The adjusted image features are fused to obtain fused image features of the target device image.
10. The image recognition method according to claim 9, characterized in that: Before the trained recognition model is used to perform convolution processing on the image features to obtain the convolved image features, the method further includes: Obtaining positive samples of device images and non-device image samples, and using the positive samples of device images to train a preset recognition model to obtain an initial trained recognition model; Using the initial training recognition model to identify the non-device image samples, and based on the recognition results, screening out non-device image samples with recognition errors from the non-device image samples to obtain device image negative samples; The initial trained recognition model is modified using the positive device image samples and the negative device image samples to obtain the trained recognition model.
11. The image recognition method according to claim 8, wherein: The step of identifying at least one device region containing device parts in the target device image based on the fused image features includes: Extracting device area features of the device area from the fused image features; determining device area information in the target device image according to the device area feature; Based on the device region information, position information of at least one device region is identified in the target device image to obtain at least one device region containing device parts in the target device image.
12. The image recognition method according to claim 8, wherein: The classifying the device area based on the fused image features includes: Extracting image features corresponding to the device area of a preset size from the fused image features to obtain target image features; Obtaining a pooling weight of the target image feature according to a size of the device area, and weighting the target image feature based on the pooling weight; The weighted image features are pooled, and the type of the corresponding device area is determined based on the pooled image features.
13. An image recognition device, characterized in that: include: A first display unit is used to display a part identification page of a client, wherein the part identification page includes an image acquisition control; a detection unit, configured to perform device area detection on the acquired target device image in response to a trigger operation on the image acquisition control, wherein the device area detection is performed based on device area identification information returned by a server, the server identifying at least one device area containing device parts in the target device image, classifying the device area, determining a type corresponding to the device area, determining part information corresponding to the type of the device area from a preset device part information set based on the type of the device area, and obtaining the device area identification information based on the part information corresponding to the type of the device area; A second display unit is configured to display an area selection page when the number of detected device areas exceeds a preset number threshold, the area selection page including a plurality of device areas to be selected; The third display unit is used to display a part identification result page in response to a selection operation on the device area, wherein the part identification result page includes a parts list of the target device area selected by the selection operation, and the parts list includes attribute information of at least one device part in the target device area.
14. An image recognition device, characterized in that: include: a receiving unit, configured to receive a target device image sent by a terminal, and extract local features of the target device image to obtain image features of multiple sizes; a fusion unit, configured to fuse the image features to obtain fused image features of the target device image; an identification unit, configured to identify at least one device region containing device parts in the target device image based on the fused image features; a classification unit, configured to classify the device area based on the fused image features, and obtain device part information corresponding to the type of the device area, thereby obtaining device area identification information of the target device image, and obtaining device part information corresponding to the type of the device area, including: determining, based on the type of the device area, part information corresponding to the type of the device area from a preset device part information set; A sending unit is configured to send the device area identification information to the terminal so that the terminal performs device area detection on the target device image based on the device area identification information.
15. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores an application program, and the processor is configured to run the application program in the memory to execute the steps of the image recognition method according to any one of claims 1 to 12.
16. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the image recognition method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Management method and apparatus for components of equipment in an electronic terminal
CN108182563A
Bridge crack detection method based on multi-resolution convolutional network
CN112348770A
Automatic generation of user interfaces using image recognition
US20200097725A1