Image acquisition and identification method and device, electronic device and storage medium
By performing high-compression RAW domain compression and low-power communication transmission on the image acquisition device, combined with a preset acquisition cycle and image recognition algorithm, the problem of making image acquisition devices lightweight and portable is solved, achieving low-power image acquisition and recognition, and improving the user experience.
Patent Information
- Application Number
- CN202411187127.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-27
- Publication Date
- 2026-03-06
AI Technical Summary
Existing image acquisition equipment is difficult to make lightweight and portable, and high-resolution image data transmission consumes a lot of energy, affecting the device's battery life and size.
By performing high-compression RAW domain compression on the raw image data acquired by the image acquisition device, and using low-power communication technology to transmit the compressed image data to a pre-bound image recognition device for recognition, combined with a preset acquisition cycle and image recognition algorithm, the power consumption of the device is reduced.
It achieves lightweight and portable image acquisition equipment, reduces power consumption and size, improves user experience, and meets image recognition needs in different scenarios.
Smart Images

Figure CN121619508A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to an image acquisition and recognition method, device, electronic device, and storage medium. Background Technology
[0002] With technological advancements, various types of cameras have emerged, allowing users to take pictures using image acquisition devices such as mobile phones and digital cameras. To meet user needs, image acquisition devices are gradually developing towards compactness and miniaturization. The application scenarios for image acquisition using cameras are becoming increasingly diverse, and correspondingly, different camera performance requirements are being placed on different usage scenarios.
[0003] It should be noted that the information disclosed in the background section above is only used to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0004] This disclosure provides an image acquisition and recognition method, device, electronic device, and storage medium.
[0005] According to a first aspect of the present disclosure, an image acquisition method is provided, comprising:
[0006] Acquire the first image data of the target scene;
[0007] The first image data is compressed to obtain the second image data;
[0008] The second image data is sent to a pre-bound image recognition device so that the image recognition device can perform image recognition on the second image data using an image recognition model.
[0009] In some embodiments of this disclosure, image compression is performed on the first image data to obtain second image data, including:
[0010] The second image data is obtained by RAW domain compression of the first image data at the target compression ratio.
[0011] In some embodiments of this disclosure, the target compression ratio is greater than or equal to 5.
[0012] In some embodiments of this disclosure, sending the second image data to a pre-bound image recognition device includes:
[0013] The second image data is transmitted to the image recognition device using short-range communication technology.
[0014] In some embodiments of this disclosure, acquiring first image data of the target scene includes:
[0015] The first image data of the target scene is collected according to a preset collection cycle;
[0016] The acquisition period is not higher than a preset period threshold.
[0017] In some embodiments of this disclosure, acquiring first image data of the target scene includes:
[0018] Acquire third image data; wherein the image resolution of the third image data is lower than that of the first image data;
[0019] Perform image recognition on the third image data;
[0020] In response to the image recognition result of the third image data matching the target scene, the first image data of the target scene is acquired.
[0021] In some embodiments of this disclosure, the first image data of the target scene includes:
[0022] In response to receiving a collection command from the image recognition device, the first image data of the target scene is collected.
[0023] In some embodiments of this disclosure, the provided image acquisition method further includes:
[0024] During non-acquisition periods, the image sensor is kept in sleep mode.
[0025] The image sensor is configured to acquire the first image data of the target scene.
[0026] According to a second aspect of the present disclosure, an image recognition method is provided, comprising:
[0027] Receive second image data corresponding to the target scene sent by a pre-bound image acquisition device;
[0028] The second image data is decoded to obtain the fourth image data;
[0029] The fourth image data is subjected to image recognition using an image recognition model to obtain the recognition information of the target scene;
[0030] Output the identification information.
[0031] In some embodiments of this disclosure, image decoding is performed on the second image data to obtain fourth image data, including:
[0032] The second image data is RAW image decoded to obtain the fourth image data.
[0033] In some embodiments of this disclosure, receiving second image data sent by a pre-bound image acquisition device includes:
[0034] The second image data is received from the pre-bound image acquisition device using short-range communication technology.
[0035] In some embodiments of this disclosure, the provided image recognition method further includes:
[0036] Receive user input information;
[0037] Specifically, image recognition is performed on the fourth image data using an image recognition model to obtain recognition information of the target scene, including:
[0038] The image recognition model performs image recognition on the fourth image data based on the user input information to obtain the recognition information of the target scene.
[0039] According to a third aspect of the present disclosure, an image acquisition device is provided, comprising:
[0040] Image sensors are used to acquire initial image data of the target scene;
[0041] An image compression unit is used to compress the first image data to obtain the second image data;
[0042] The microcontroller unit is used to send the second image data to a pre-bound image recognition device, so that the image recognition device can perform image recognition on the second image data through an image recognition model.
[0043] According to a fourth aspect of the present disclosure, an image recognition device is provided, comprising:
[0044] An image receiving unit is used to receive second image data sent by a pre-bound image acquisition device; wherein the second image data is obtained by the image acquisition device compressing first image data of the target scene;
[0045] An image decoding unit is used to decode the second image data to obtain the fourth image data;
[0046] An image recognition unit is used to perform image recognition on the fourth image data through an image recognition model to obtain recognition information of the target scene;
[0047] The output unit is used to output the identification information.
[0048] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising:
[0049] processor;
[0050] Memory used to store processor-executable instructions;
[0051] The processor is configured to implement the image acquisition method described in the first aspect above.
[0052] According to a sixth aspect of the present disclosure, an electronic device is provided, comprising:
[0053] processor;
[0054] Memory used to store processor-executable instructions;
[0055] The processor is configured to implement the image recognition method described in the first aspect above.
[0056] According to a seventh aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, which, when the instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform any of the image acquisition methods described in the first aspect.
[0057] According to an eighth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, which, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform any of the image recognition methods described in the second aspect above.
[0058] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects:
[0059] This disclosure involves acquiring first image data of a target scene using an image acquisition device; compressing the first image data to obtain second image data; and sending the second image data to a pre-bound image recognition device, which then performs image recognition on the second image data using an image recognition model. By using the image recognition model for image recognition, scene recognition can be performed on the acquired images to meet users' needs for scene recognition and provide a better user experience.
[0060] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0061] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0062] Figure 1 This is a flowchart illustrating an image acquisition method according to some embodiments of the present disclosure.
[0063] Figure 2 This is an implementation process flow of step S102 shown in some embodiments of this disclosure. Figure 1 .
[0064] Figure 3 This is an implementation process flow of step S102 shown in some embodiments of this disclosure. Figure 2 .
[0065] Figure 4 This is an implementation process flow of step S102 shown in some embodiments of this disclosure. Figure 3 .
[0066] Figure 5 This is a flowchart of an image recognition method according to some embodiments of the present disclosure. Figure 1 .
[0067] Figure 6 This is a flowchart of an image recognition method according to some embodiments of the present disclosure. Figure 2 .
[0068] Figure 7 This is a schematic diagram of the architecture of an AI recognition system constructed according to an exemplary embodiment of the present disclosure.
[0069] Figure 8 This is illustrated according to an exemplary embodiment. Figure 7 The diagram shows an application scenario of the AI recognition system.
[0070] Figure 9 This is a block diagram illustrating an image acquisition device according to an exemplary embodiment of the present disclosure.
[0071] Figure 10 This is a frame of an image recognition device shown according to some embodiments of the present disclosure. Figure 2 .
[0072] Figure 11 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. Detailed Implementation
[0073] Some embodiments of this disclosure will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. Various changes, modifications, and equivalents of the methods, apparatus, and / or systems described herein will become apparent upon understanding this disclosure. For example, the order of operations described herein is merely illustrative and is not limited to those orders set forth herein, but can be changed as will become apparent upon understanding this disclosure, except for operations that must be performed in a particular order. Furthermore, for clarity and brevity, descriptions of features known in the art may be omitted.
[0074] The embodiments described in the following examples of this disclosure are not representative of all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0075] The specific implementation methods of the embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.
[0076] Figure 1 This is a flowchart illustrating an image acquisition method according to some embodiments of the present disclosure. Figure 1 ,like Figure 1 As shown, the image acquisition method can be applied to electronic devices capable of image acquisition, including but not limited to smartphones, wearable devices with cameras, cameras, and smart tablets.
[0077] Figure 1 The image acquisition method shown includes the following steps.
[0078] In step S102, the first image data of the target scene is acquired.
[0079] It should be noted that in some embodiments of this disclosure, Figure 1 The image acquisition method shown is applied to an image acquisition device. The image acquisition device can acquire image data from different scenes. In conjunction with an image recognition device pre-bonded to the image acquisition device, the image acquisition device acquires image data for a target scene.
[0080] In exemplary embodiments of this disclosure, the target scene can be a scene where the image recognition device intends to capture an image. For example, when using AR (Augmented Reality) street view navigation, the navigation device needs to capture street view images and perform image recognition to extract street view information for navigation. In this example, the street view to be captured is the target scene. The target scene can also be a scene where the user intends to capture an image. For example, if the user wants to know information about a vase in front of them, an image acquisition device can capture an image of the vase, and an image recognition device can identify it to obtain the information the user wants to know about the vase. In this example, the vase-capturing scene is the target scene. Those skilled in the art will understand that the above examples are merely illustrative and are not intended to limit the scope of the target scene in the embodiments of this disclosure.
[0081] In an exemplary embodiment of this disclosure, the first image data is a RAW format image, which is an unprocessed raw image data format that contains all the raw information captured by the sensor, including data such as color depth, exposure, and white balance.
[0082] In an exemplary embodiment of this disclosure, in order to ensure accuracy during subsequent image processing and to have sufficient image information for image recognition, the resolution of the first image data is higher than a preset resolution threshold. For example, the preset resolution threshold may be 720p, or the resolution of the first image data may be 1080p.
[0083] In step S104, the first image data is compressed to obtain the second image data.
[0084] It should be noted that the first image data is image data generated by the image sensor acquiring the target scene, and the data volume is relatively large. In the exemplary embodiment of this disclosure, the first image data is a RAW format image, and the data size of a single RAW image may be as high as tens of megabytes. If the original RAW format image is transmitted, not only will the image acquisition device consume more resources and storage space when processing the original RAW format image, but the amount of data transmitted from the image acquisition device to the image recognition device will also be large, requiring more energy for transmission, which is not conducive to the low power consumption of the image acquisition device. Therefore, the first image data is compressed to obtain the second image data.
[0085] In an exemplary embodiment of this disclosure, step S104 specifically involves performing RAW domain compression on the first image data at a target compression ratio to obtain the second image data, significantly reducing the data volume while retaining sufficient image quality. To achieve a higher compression ratio and minimize the data volume of the second image data, RAW domain compression can employ lossy compression algorithms or hybrid compression algorithms. Lossy compression algorithms sacrifice some image quality to achieve a higher compression ratio. This can be achieved through transform coding, such as Discrete Cosine Transform (DCT) or Wavelet Transform, to remove redundant information in the image. Predictive coding can also be used to predict the relationship between adjacent pixels to reduce the data volume. Hybrid compression algorithms combine the advantages of lossless and lossy compression techniques, retaining key information while reducing the data volume.
[0086] In some embodiments of this disclosure, the target compression ratio is greater than or equal to 5, for example, the target compression ratio can be 10, 12, or 15. Those skilled in the art will understand that the above-mentioned target compression ratio values are merely examples and are not intended to limit the scope of protection of the embodiments of this disclosure. The target compression ratio is sufficient to ensure that the amount of compressed second image data meets the transmission requirements and can be set according to actual needs.
[0087] In step S106, the second image data is sent to the pre-bound image recognition device so that the image recognition device can perform image recognition on the second image data through the image recognition model.
[0088] In the embodiments of this disclosure, the image acquisition device and the image recognition device are pre-bonded, for example, a pre-bonded AI camera and a mobile phone, with the AI camera and mobile phone having pre-established Bluetooth communication. Specifically, short-range communication technology is used to transmit the second image data to the image recognition device. The image acquisition device can use Bluetooth communication technology to wirelessly transmit the second image data to the image recognition device, reducing the power consumption of the image acquisition device by leveraging the low-power characteristics of Bluetooth transmission.
[0089] In some exemplary embodiments of this disclosure, Bluetooth Low Energy (BLE) communication technology can be used to transmit second image data to an image recognition device. BLE is designed specifically for low-power applications, allowing devices to consume very little energy while transmitting data, making it ideal for small, battery-powered devices requiring long-term operation. Furthermore, BLE supports rapid connection establishment and termination to reduce unnecessary power consumption. However, BLE's low bandwidth results in a lower data transmission rate; therefore, to accommodate the low bandwidth, the amount of second image data cannot be too large, thus setting the target compression ratio.
[0090] The applicant discovered that, in order to ensure the realization of the image acquisition function and to ensure that the device can operate for a period of time, the image acquisition device in the relevant technology needs to be equipped with a power supply module that can provide sufficient energy. However, the size and weight of this power supply module cannot be reduced further, and each component still needs to maintain a certain size and weight. This results in the existing image acquisition devices being relatively large in weight and size, making it difficult to achieve lightweight and portable design.
[0091] As can be seen, in order to reduce device power consumption and achieve lightweight image acquisition equipment, the applicant proposes an image acquisition method provided in this disclosure embodiment. This method compresses first image data of the target scene and then sends the compressed second image data to a pre-bound image recognition device. The image recognition device then uses an image recognition model to perform image recognition on the second image data. Based on the characteristic that image recognition models do not have high requirements for the accuracy and information content of input image data, images can be compressed at a relatively high target compression ratio. Furthermore, low-power communication technology can be used to transmit the second image data to the image recognition device, thereby significantly reducing the operating power consumption of the image acquisition equipment and achieving lightweight and portable design.
[0092] As can be seen from the above steps, the image acquisition method applied to an image acquisition device provided in the embodiments of this disclosure acquires first image data of a target scene; compresses the first image data to obtain second image data; and sends the second image data to a pre-bound image recognition device, so that the image recognition device can perform image recognition on the second image data using an image recognition model. By cooperating with the image recognition device and the image recognition model, scene recognition can be performed on the acquired images to meet the user's needs for scene recognition and provide a better user experience.
[0093] Furthermore, since the image recognition model does not have high requirements for image clarity, the image acquisition method provided in the embodiments of this disclosure for image acquisition devices can compress the image after acquisition to obtain second image data, thereby greatly reducing the resources required for the operation and data transmission of the image acquisition device, thus reducing the power consumption of the image acquisition device. Therefore, a smaller size and lower capacity power supply module can be used, thereby providing a low-power and lightweight image acquisition device.
[0094] like Figure 2 The following is a flowchart illustrating the implementation process of step S102 provided for some exemplary embodiments of this disclosure. Figure 1 The process includes the following steps.
[0095] In step S202, the first image data of the target scene is acquired according to a preset acquisition cycle.
[0096] It should be noted that, in order to further reduce the power consumption of the image acquisition device, the first image data can be acquired according to a preset acquisition cycle, and the acquisition cycle is not higher than the preset cycle threshold. Those skilled in the art will understand that the preset threshold is set considering the actual shooting needs and low power consumption requirements, and this disclosure embodiment does not limit it. For example, the preset cycle threshold can be 2 seconds, and correspondingly, the acquisition cycle can be 1 second, that is, the image acquisition device acquires the first image data of the target scene once every 1 second to obtain a first image.
[0097] In some exemplary embodiments of this disclosure, the image acquisition device can convert optical images into electrical signals for image acquisition by setting an image sensor; that is, the image sensor is configured to acquire first image data of the target scene. In specific implementations, a low-power, low-cost CMOS sensor can be used to ensure low operating power consumption of the image acquisition device.
[0098] This disclosure reduces the frequency of image sensor acquisition by acquiring first image data of the target scene according to a preset acquisition cycle, rather than capturing the target scene in real time, thereby reducing the number of image acquisitions and the amount of data, and thus reducing the operating power consumption of the image acquisition device.
[0099] like Figure 3 The following is a flowchart illustrating the implementation process of step S102 provided for some exemplary embodiments of this disclosure. Figure 2 The process includes the following steps.
[0100] In step S302, third image data is acquired.
[0101] It should be noted that the resolution of the third image data is lower than that of the first image data. In practice, the third image data can be acquired periodically or continuously. There are no restrictions on the image acquisition scenario. Image acquisition devices are used for image acquisition, but the resolution of the acquired images is not high, lower than that of the first image data. Since the acquired third image data is not transmitted to an image recognition device for image recognition, the requirements for the third image data are lower than those for the first image data. Therefore, to further reduce unnecessary device power consumption, the resolution of the acquired third image data is set to be low.
[0102] In step S304, image recognition is performed on the third image data.
[0103] It should be noted that the image acquisition device can use image recognition algorithms to perform image recognition on the third-party image data to determine whether the image recognition result of the third-party image data matches the target scene. For example, if the target scene is a food scene, the image acquisition device can be controlled to collect third-party image data after the user enters the dining area. It can be understood that the image recognition result of the collected third-party image data includes the environmental scene, the table and chair scene, the people scene, and the food scene.
[0104] In step S306, in response to the image recognition result of the third image data matching the target scene, the first image data of the target scene is acquired.
[0105] In this embodiment of the disclosure, in response to the image recognition result of the third image data matching the target scene, the image acquisition device acquires the first image data again to ensure that while ensuring that the image acquisition device can acquire images of the target scene, the number of high-resolution images acquired is also reduced, thereby reducing the resources consumed during image acquisition and reducing the power consumption of the device.
[0106] In some exemplary embodiments of this disclosure, the image acquisition device can acquire images by setting an image sensor, that is, the image sensor is configured to acquire first image data and third image data. A Neural Processing Unit (NPU) is set to perform image recognition on the third image data. After determining that the image recognition result of the third image data matches the target scene, an instruction is issued to the image sensor to acquire the first image data. It should be noted that an NPU is a chip specifically designed for neural network calculations. NPU chips typically feature high integration and low power consumption, enabling devices to execute artificial intelligence (AI) algorithms faster and improving the device's AI computing capabilities.
[0107] This disclosure pre-collects third image data, performs image recognition on the third image data, determines that the image recognition result matches the target scene, and collects first image data. This ensures that while the image acquisition device can acquire images of the target scene, the number of high-resolution images acquired is also reduced, thereby reducing the resources consumed during image acquisition and the amount of data transmitted in the images, thus reducing the power consumption of the device.
[0108] like Figure 4 The following is a flowchart illustrating the implementation process of step S102 provided for some exemplary embodiments of this disclosure. Figure 3 The process includes the following steps.
[0109] In step S402, in response to receiving the acquisition command sent by the image recognition device, the first image data of the target scene is acquired.
[0110] In some embodiments of this disclosure, the image recognition device can be a smart terminal such as a mobile phone. When a user uses a smart terminal, an image acquisition device is required to cooperate in image acquisition. At this time, the image recognition device sends an acquisition command to the image acquisition device so that the image acquisition device can acquire the first image data of the target scene.
[0111] For example, when a user uses AR street view navigation on their mobile phone, the image acquisition device is attached to the user's glasses frame in the form of a miniature camera. The mobile phone needs to acquire street view images. On the one hand, it can rely on the mobile phone's own camera to acquire images. On the other hand, the mobile phone sends an acquisition command to the image acquisition device. After receiving the acquisition command, the image acquisition device parses the acquisition command to determine that the target scene is a street view, and controls the image sensor set in the image acquisition device to acquire street view images.
[0112] This disclosure acquires first image data of a target scene in response to receiving an acquisition command sent by an image recognition device, thereby establishing a connection between the image acquisition device and the image recognition device, achieving intelligent recognition, better serving users, and improving user experience.
[0113] In this embodiment of the disclosure, an image acquisition method is provided, which further includes controlling the image sensor to be in a sleep state during non-acquisition periods. It should be noted that the image sensor is configured to acquire first image data of the target scene, that is, an image sensor is set up within the image acquisition device to perform image acquisition. During non-image acquisition periods, the image sensor can be controlled to be in a sleep state to avoid the image sensor being constantly in working state and to reduce unnecessary power consumption of the image acquisition device.
[0114] In an exemplary embodiment of this disclosure, the image sensor can be configured to periodically identify the scene and acquire image data at an image frame rate of 1 FPS, that is, to capture one frame per second. The exposure time of the image sensor is set to a short exposure time, and the AE / AWB (automatic exposure / automatic white balance) function of the image sensor is enabled to minimize motion blur in the acquired images. In other words, the image sensor captures images only for a very short period within a 1-second acquisition cycle, for example, one frame can be captured in 33ms. During the remaining time period, the image sensor is in a sleep state, thereby reducing the power consumption required by the image sensor and ensuring low power consumption of the image acquisition device.
[0115] This disclosure embodiment achieves low power consumption in the image acquisition device, enabling the use of a smaller, lower-capacity power supply unit to reduce the weight and volume of the power supply unit. This significantly reduces the weight and volume of the image acquisition device, achieving lightweight design. In specific examples, the image acquisition device can weigh less than 5g, making it more portable and convenient for users to carry. By cooperating with pre-attached image recognition devices such as mobile phones, the image acquisition device can provide users with more intelligent services, improve user experience, and align with future trends in intelligent development.
[0116] Figure 5 This is a flowchart of an image recognition method according to some embodiments of the present disclosure. Figure 1 ,like Figure 5 As shown, the image recognition method can be applied to electronic devices, including but not limited to terminal devices such as smartphones, wearable devices, and smart tablets, and can also include server-side devices such as local servers and cloud servers, which can be deployed in a computer cluster consisting of one computer or multiple computers.
[0117] Figure 5 The image recognition method shown includes the following steps.
[0118] In step S502, second image data corresponding to the target scene is received from a pre-bound image acquisition device.
[0119] It should be noted that in some embodiments of this disclosure, such as Figure 5 The image recognition method shown is applied to an image recognition device. The image recognition device is pre-bonded with an image acquisition device, such as a pre-bonded AI camera and a mobile phone, which have pre-established Bluetooth communication.
[0120] In some embodiments of this disclosure, short-range communication technology is used to receive second image data sent by a pre-bound image acquisition device. Specifically, Bluetooth communication technology is used to wirelessly receive the second image data, enabling the formation of a communication network for image data transmission without the need for third-party devices such as routers.
[0121] In step S504, the second image data is decoded to obtain the fourth image data.
[0122] In an exemplary embodiment of this disclosure, the second image data is RAW image decoding to obtain the fourth image data. To reduce the amount of data transmitted, the image acquisition device compresses the image data using RAW domain compression at a high target compression ratio. When decoding the second image data, the image recognition device employs a RAW image decoding algorithm that matches the RAW domain compression to decode the second image data into the fourth image data.
[0123] In step S506, the fourth image data is subjected to image recognition by an image recognition model to obtain the recognition information of the target scene.
[0124] In this embodiment of the disclosure, the image recognition model is a large AI model that recognizes input image data. It is a large neural network model trained using deep learning and other technologies. It typically has a large number of parameters and a complex architecture, and can achieve high-precision image recognition tasks.
[0125] In an exemplary embodiment of this disclosure, the image recognition model can be a large intelligent recognition model deployed locally on the image recognition device, or it can be a large model deployed in the cloud, such as a multimodal large model. By calling the cloud service of the multimodal large model, the image data can be recognized in a timely manner, so as to ensure that the image data can be recognized in a timely manner even when the processing capability of the model deployed on the image recognition device is insufficient, and the recognition result can be fed back to the user.
[0126] In an exemplary embodiment of this disclosure, the decoded fourth image data may contain a lot of irrelevant information or not conform to the format requirements of the image recognition model. The image recognition device may also input the fourth image data into the image signal processor (ISP) set in the image recognition device to perform various processing operations such as noise suppression, white balance adjustment, color correction, and sharpening, so as to improve the recognition accuracy of the subsequent image recognition model, reduce the computing power consumption of the image recognition model, improve recognition efficiency, and enhance the user experience.
[0127] In step S508, the identification information is output.
[0128] In some embodiments of this disclosure, the image recognition device outputs recognition information to the user to serve the user. Specifically, the information can be displayed visually to provide feedback to the user, or the recognition information can be converted from text to speech, and the converted speech data can be notified to the user through a device such as headphones paired with the image recognition device to further enhance the user experience.
[0129] As can be seen from the above steps, this disclosure receives second image data corresponding to the target scene sent by a pre-bound image acquisition device, decodes the second image data to obtain fourth image data, performs image recognition on the fourth image data using an image recognition model, obtains the recognition information of the target scene, and outputs it. This enables the matching recognition of compressed image data acquired by the image acquisition device and provides feedback recognition information to the user, thus achieving the cooperative use of the image recognition device and the image acquisition device to provide intelligent services to the user.
[0130] In this disclosure, an image recognition method is provided, such as... Figure 6 The diagram shown is a flowchart of an image recognition method according to an exemplary embodiment of this disclosure. Figure 2 .
[0131] In this embodiment of the disclosure, Figure 6 In the image recognition method shown, steps S602, S604, and S610 are respectively related to... Figure 5 Steps S502, S504, and S508 in the image recognition method shown correspond to each other and will not be repeated here.
[0132] exist Figure 5 Based on the image recognition method shown, Figure 6 The image recognition method shown may also include the following steps.
[0133] In step S606, user input information is received.
[0134] It's important to note that user input represents the user's needs. For example, a user might want to use their phone to identify information about a garment, such as its price, purchase platform, fabric, and style. The user could access an interactive smart model tool on their phone, such as ChatGPT, and input: "Identify this garment and provide information such as its price, purchase platform, fabric, and style."
[0135] Accordingly, step S506 is changed to step S608.
[0136] In step S608, the fourth image data is image recognized based on the user input information using an image recognition model to obtain the recognition information of the target scene.
[0137] In the above exemplary implementation, in response to this input information, the mobile phone sends a capture command to the image acquisition device, so that the image acquisition device can capture an image of the clothing. After receiving the clothing image, the mobile phone inputs the clothing image and the user's input request to identify the clothing in front of it and provide relevant information such as price, purchase platform, fabric, and style into ChatGPT. ChatGPT understands the user's needs based on the user's input information and combines the understood user needs to recognize the clothing image to obtain recognition information to meet the user's needs.
[0138] Figure 7 This is an architecture diagram of an AI recognition system constructed using the image acquisition method and image recognition method provided in the embodiments of this disclosure, according to an exemplary embodiment of this disclosure. Figure 7 As can be seen, the AI recognition system comprises two parts: an AI camera that applies the image acquisition method provided in the embodiments of this disclosure, and a smartphone that applies the image recognition method provided in the embodiments of this disclosure.
[0139] AI cameras include an image sensor, a RAW compression chip, a low-power microcontroller unit (MCU), a Bluetooth transceiver unit, and a power supply unit.
[0140] The image sensor is equipped with a DVP / 8 (Digital Video Port) and a MIPI (Mobile Industry Processor Interface) to transmit image data output from the sensor. The image sensor sends the acquired first image data to a RAW compression chip via the DVP / 8 for RAW domain compression. The compression ratio is no less than 10, achieving near-lossless visual quality after compression. The RAW compression chip is a hardware compression circuit that implements the RAW domain compression algorithm. The RAW compression chip then sends the compressed second image data to an SPI (Serial Peripheral Interface) on the MCU. After receiving the second image data, the MCU transmits the data to the smartphone via a Bluetooth transceiver unit.
[0141] The image sensor sends the acquired third image data to the MCU via MIPI. The MCU uses the configured NPU to perform image recognition on the third image data. When it is determined that the image recognition result matches the target scene, the NPU issues the first image data acquisition command, which is sent to the image sensor via the SDI (Serial Digital Interface) configured in the MCU to control the image sensor to perform image acquisition.
[0142] After receiving the acquisition command sent by the smartphone, the Bluetooth transceiver unit of the AI camera analyzes the acquisition command and generates an image acquisition command. This command is then sent to the image sensor via the SDI set in the MCU to control the image acquisition device to acquire the first image data of the target scene.
[0143] The power supply units are for the Sensor, RAW compression chip, MCU power supply and Bluetooth transceiver unit, to provide the energy required for operation.
[0144] The smartphone-side architecture includes a Bluetooth transceiver unit, memory, a RAW decoding circuit, an ISP, and a large-scale model. After receiving the second image data from the AI camera, the Bluetooth transceiver unit first stores it in memory. The RAW decoding circuit retrieves the second image data from memory and performs RAW image decoding to obtain the fourth image data. This fourth image data is then input into the ISP for image processing. The processed image data is then input into the large-scale model for image recognition. The large-scale model can also receive user input. When the processing power of the large-scale model is insufficient, the phone uses a network connection to input the processed image data into the cloud-based multimodal large-scale model Nass service system, leveraging the computing power of the cloud-based model for image recognition. User input can also be uploaded to the cloud-based large-scale model.
[0145] Figure 8 This is a schematic diagram illustrating the application scenarios of the aforementioned AI recognition system.
[0146] like Figure 8 As shown, in this application scenario, the AI camera 801 captures images of the target scene and transmits the compressed image data to the smartphone 802 via Bluetooth. After receiving the data, the smartphone 802 decodes the compressed image data and uses a large model for recognition to obtain the recognition information required by the user. It then converts the data into voice data and transmits it to the Bluetooth headset 803 worn by the user via Bluetooth communication to notify the user.
[0147] Since the AI camera weighs no more than 5g, it can be magnetically attached to the temple of the glasses near the hinge without affecting the balance of the glasses. The AI camera's casing has a magnetic structure that allows it to be placed on the temple of the glasses near the hinge.
[0148] It should be noted that the acquisition, storage, use, and processing of data in this disclosed technical solution all comply with the relevant provisions of national laws and regulations.
[0149] The following are embodiments of the apparatus disclosed herein, which can be used to execute embodiments of the method disclosed herein. For details not disclosed in the apparatus embodiments of this disclosure, please refer to the embodiments of the method disclosed herein.
[0150] Figure 9 This is a block diagram of an image acquisition device according to some exemplary embodiments of the present disclosure. (Refer to...) Figure 9 The device includes an image sensor 901, an image compression unit 902, and a microcontroller unit 903.
[0151] Image sensor 901 is used to acquire first image data of the target scene;
[0152] The image compression unit 902 is used to compress the first image data to obtain the second image data;
[0153] The microcontroller unit 903 is used to send the second image data to a pre-bound image recognition device so that the image recognition device can perform image recognition on the second image data through an image recognition model.
[0154] It should be noted that the device in this embodiment can be applied to image acquisition equipment.
[0155] In some exemplary embodiments of this disclosure, the image sensor 901 is configured to acquire first image data of a target scene according to a preset acquisition period; wherein the acquisition period is not higher than a preset period threshold.
[0156] In some exemplary embodiments of this disclosure, the image sensor 901 is configured to: acquire third image data; wherein the image resolution of the third image data is lower than that of the first image data; perform image recognition on the third image data; and acquire first image data of the target scene in response to the image recognition result of the third image data matching the target scene.
[0157] In some exemplary embodiments of this disclosure, the image sensor 901 is configured to: acquire first image data of a target scene in response to receiving an acquisition command sent by an image recognition device.
[0158] In some exemplary embodiments of this disclosure, the image compression unit 902 is configured to: perform RAW domain compression on the first image data at a target compression ratio to obtain the second image data.
[0159] In some exemplary embodiments of this disclosure, the target compression ratio is greater than or equal to 5.
[0160] In some exemplary embodiments of this disclosure, the microcontroller unit 903 is configured to transmit second image data to an image recognition device using short-range communication technology.
[0161] In some exemplary embodiments of this disclosure, the microcontroller unit 903 is also configured to control the image sensor 901 to be in a sleep state during non-acquisition periods.
[0162] Figure 10 This is a block diagram illustrating an image recognition device according to some exemplary embodiments of the present disclosure. (Refer to...) Figure 10 The device includes an image receiving unit 1001, an image decoding unit 1002, an image recognition unit 1003, and an output unit 1004.
[0163] The image receiving unit 1001 is used to receive second image data sent by a pre-bound image acquisition device; wherein, the second image data is obtained by the image acquisition device compressing the first image data of the target scene;
[0164] The image decoding unit 1002 is used to decode the second image data to obtain the fourth image data;
[0165] The image recognition unit 1003 is used to perform image recognition on the fourth image data through an image recognition model to obtain recognition information of the target scene;
[0166] Output unit 1004 is used to output identification information.
[0167] It should be noted that the device in this embodiment can be applied to image recognition devices.
[0168] In some exemplary embodiments of this disclosure, the image receiving unit 1001 is configured to receive second image data sent by a pre-bound image acquisition device using short-range communication technology.
[0169] In some exemplary embodiments of this disclosure, the image decoding unit 1002 is configured to perform RAW image decoding on the second image data to obtain the fourth image data.
[0170] In some exemplary embodiments of this disclosure, the image recognition apparatus further includes: an information receiving unit for receiving user input information. Accordingly, the image recognition unit 1003 is configured to: perform image recognition on the fourth image data based on the user input information using an image recognition model to obtain recognition information of the target scene.
[0171] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0172] Figure 11 This is a block diagram illustrating an electronic device according to an exemplary embodiment of the present disclosure. For example, device 1100 may be a mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness device, personal digital assistant, etc.
[0173] Reference Figure 11The device 1100 may include one or more of the following components: a processing component 1102, a memory 1104, a power supply component 1106, a multimedia component 1108, an audio component 1110, an input / output (I / O) interface 1112, a sensor component 1114, and a communication component 1116.
[0174] Processing component 1102 typically controls the overall operation of device 1100, such as operations associated with display, telephone calls, data communication, camera operation, and recording operations. Processing component 1102 may include one or more processors 1120 to execute instructions to perform all or part of the steps of the methods described above. Furthermore, processing component 1102 may include one or more modules to facilitate interaction between processing component 1102 and other components. For example, processing component 1102 may include a multimedia module to facilitate interaction between multimedia component 1108 and processing component 1102.
[0175] Memory 1104 is configured to store various types of data to support the operation of device 1100. Examples of such data include instructions for any application or method operating on device 1100, contact data, phonebook data, messages, pictures, videos, etc. Memory 1104 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0176] Power supply component 1106 provides power to various components of device 1100. Power supply component 1106 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to device 1100.
[0177] Multimedia component 1108 includes a screen that provides an output interface between the device 1100 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors may sense not only the boundaries of the touch or swipe action but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 1108 includes a front-facing camera and / or a rear-facing camera. When the device 1100 is in an operating mode, such as a shooting mode or a video mode, the front-facing camera and / or the rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0178] Audio component 1110 is configured to output and / or input audio signals. For example, audio component 1110 includes a microphone (MIC) configured to receive external audio signals when device 1100 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 1104 or transmitted via communication component 1116. In some embodiments, audio component 1110 also includes a speaker for outputting audio signals.
[0179] I / O interface 1112 provides an interface between processing component 1102 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, power buttons, and lock buttons.
[0180] Sensor assembly 1114 includes one or more sensors for providing status assessments of various aspects of device 1100. For example, sensor assembly 1114 may detect the on / off state of device 1100, the relative positioning of components such as the display and keypad of device 1100, changes in the position of device 1100 or a component of device 1100, the presence or absence of user contact with device 1100, the orientation or acceleration / deceleration of device 1100, and temperature changes of device 1100. Sensor assembly 1114 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 1114 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, sensor assembly 1114 may also include an accelerometer, a gyroscope, a magnetometer, a pressure sensor, or a temperature sensor.
[0181] Communication component 1116 is configured to facilitate wired or wireless communication between device 1100 and other devices. Device 1100 can access wireless networks based on communication standards, such as WiFi, 3G, 4G, 5G, other communication standards, or combinations thereof. In some embodiments of this disclosure, communication component 1116 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In some embodiments of this disclosure, communication component 1116 further includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on radio frequency identification (RFID) technology, Infrared Data Association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0182] In some embodiments of this disclosure, the apparatus 1100 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the methods described above.
[0183] In some embodiments of this disclosure, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1104 including instructions, which can be executed by a processor 1120 of device 1100 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0184] In some embodiments of this disclosure, a non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enables the mobile terminal to perform an image acquisition method or an image recognition method, the image acquisition method comprising:
[0185] First image data of the target scene is acquired; second image data is obtained by image compression of the first image data; the second image data is sent to a pre-bound image recognition device so that the image recognition device can perform image recognition on the second image data through an image recognition model.
[0186] The image recognition method includes:
[0187] Receive second image data corresponding to the target scene sent by a pre-bound image acquisition device; decode the second image data to obtain fourth image data; perform image recognition on the fourth image data through an image recognition model to obtain recognition information of the target scene; output the recognition information.
[0188] In some embodiments of this disclosure, a computer program product is also provided, including a computer program / instructions, which, when executed by a processor, implement an image acquisition method or an image recognition method, wherein the image acquisition method includes:
[0189] First image data of the target scene is acquired; second image data is obtained by image compression of the first image data; the second image data is sent to a pre-bound image recognition device so that the image recognition device can perform image recognition on the second image data through an image recognition model.
[0190] The image recognition method includes:
[0191] Receive second image data corresponding to the target scene sent by a pre-bound image acquisition device; decode the second image data to obtain fourth image data; perform image recognition on the fourth image data through an image recognition model to obtain recognition information of the target scene; output the recognition information.
[0192] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0193] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image acquisition method, characterized in that, The method comprises: collecting first image data of a target scene; performing image compression on the first image data to obtain second image data; sending the second image data to a pre-bound image recognition device, so that the image recognition device performs image recognition on the second image data through an image recognition model.
2. The image acquisition method of claim 1, wherein, The image compression on the first image data to obtain second image data comprises: performing RAW domain compression on the first image data at a target compression ratio to obtain the second image data.
3. The image acquisition method of claim 2, wherein, The target compression ratio is greater than or equal to 5.
4. The image acquisition method of claim 1, wherein, The sending of the second image data to the pre-bound image recognition device comprises: transmitting the second image data to the image recognition device by using short-distance communication technology.
5. The image acquisition method of claim 1, wherein, The collecting of the first image data of the target scene comprises: collecting the first image data of the target scene according to a preset collection period; wherein the collection period is not higher than a preset period threshold.
6. The image acquisition method of claim 1, wherein, The collecting of the first image data of the target scene comprises: collecting third image data, wherein the image resolution of the third image data is lower than that of the first image data; performing image recognition on the third image data; in response to the image recognition result of the third image data being consistent with the target scene, collecting the first image data of the target scene.
7. The image acquisition method of claim 1, wherein, The collecting of the first image data of the target scene comprises: in response to receiving a collection instruction sent by the image recognition device, collecting the first image data of the target scene.
8. The image acquisition method according to any one of claims 5 to 7, characterized in that, The method further comprises: controlling the image sensor to be in a sleep state during a non-collection period; wherein the image sensor is configured to collect the first image data of the target scene.
9. An image recognition method characterized by, The method comprises: receiving second image data corresponding to a target scene sent by a pre-bound image collection device; performing image decoding on the second image data to obtain fourth image data; performing image recognition on the fourth image data through an image recognition model to obtain recognition information of the target scene; outputting the recognition information.
10. The image recognition method of claim 9, wherein, The image decoding on the second image data to obtain fourth image data comprises: performing RAW image decoding on the second image data to obtain the fourth image data.
11. The image recognition method of claim 9, wherein, The receiving of the second image data sent by the pre-bound image collection device comprises: receiving the second image data sent by the pre-bound image collection device by using short-distance communication technology.
12. The image recognition method of claim 9, wherein, The method further comprises: receiving user input information; wherein the image recognition on the fourth image data through the image recognition model to obtain the recognition information of the target scene comprises: performing image recognition on the fourth image data through the image recognition model based on the user input information to obtain the recognition information of the target scene.
13. An image capture device, comprising: The method comprises: an image sensor configured to collect first image data of a target scene; an image compression unit configured to perform image compression on the first image data to obtain second image data; a microcontroller unit configured to send the second image data to a pre-bound image recognition device, so that the image recognition device performs image recognition on the second image data through an image recognition model.
14. An image recognition device, characterized by, The method comprises: An image receiving unit is configured to receive second image data sent by a pre-bound image acquisition device; wherein the second image data is obtained by image compression on first image data of a target scene by the image acquisition device; An image decoding unit is configured to perform image decoding on the second image data to obtain fourth image data; An image recognition unit is configured to perform image recognition on the fourth image data by an image recognition model to obtain recognition information of the target scene; An output unit is configured to output the recognition information.
15. An electronic device, comprising: comprise: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to implement the image acquisition method of any one of claims 1 to 8 or the image recognition method of any one of claims 9 to 12.
16. A non-transitory computer-readable storage medium, when instructions in the storage medium are executed by a processor of a mobile terminal, enabling the mobile terminal to perform the image acquisition method of any one of claims 1 to 8 or the image recognition method of any one of claims 9 to 12.