Side-by-side image detection method and electronic device using the method
By using a convolutional neural network model to identify side-by-side image formats, the problem of inaccurate identification of side-by-side image formats in existing technologies is solved, thereby improving the display effect and user experience of 3D displays.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ACER INC
- Filing Date
- 2022-02-11
- Publication Date
- 2026-07-17
AI Technical Summary
Existing technologies struggle to accurately identify image content in side-by-side image formats, causing 3D displays to fail to display images correctly.
A convolutional neural network model is used to detect whether an image conforms to the side-by-side image format. The model is trained using machine learning techniques to identify the image content in the side-by-side image format.
It improves the user experience and application scope of 3D display technology, and ensures that 3D displays can correctly display image content in side-by-side image format.
Smart Images

Figure CN115187503B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an electronic device, and more particularly to a side-by-side image detection method and an electronic device using the method. Background Technology
[0002] With advancements in display technology, displays supporting three-dimensional (3D) image playback have become increasingly common. The difference between 3D and two-dimensional (2D) displays lies in the fact that 3D technology allows viewers to experience a sense of depth in images, such as three-dimensional facial features and depth of field, an effect that traditional 2D images cannot achieve. The principle of 3D display technology is to allow the viewer to see the image with their left eye and the image with their right eye, thus creating a 3D visual effect. With the rapid development of 3D stereoscopic display technology, it can provide people with an immersive visual experience. It is known that 3D displays require specific 3D image formats and corresponding 3D display technologies for playback; otherwise, the display will fail to display the image correctly. Therefore, accurately identifying image content conforming to specific 3D image formats is a topic of concern for those skilled in the art. Summary of the Invention
[0003] This disclosure relates to a side-by-side image detection method and an electronic device using the method, which can accurately detect image content conforming to the side-by-side image format.
[0004] This invention provides a side-by-side image detection method, comprising the following steps: acquiring a first image having a first image size; and detecting a second image in the first image that conforms to a side-by-side image format, wherein the second image has a second image size, using a convolutional neural network model.
[0005] This invention provides an electronic device including a storage device and a processor. The processor is connected to the storage device and configured to perform the following steps: acquiring a first image having a first image size; detecting a second image in the first image that conforms to a side-by-side image format, wherein the second image has a second image size, using a convolutional neural network model.
[0006] Based on the above, in this embodiment of the invention, a convolutional neural network model from the field of machine learning can be used to accurately detect whether an image includes image content conforming to a side-by-side image format. This detection result can be used in various application scenarios, thereby improving the user experience and application scope of 3D display technology. Attached Figure Description
[0007] Figure 1A and Figure 1BThis is a schematic diagram of an electronic device according to an embodiment of the present invention;
[0008] Figure 2 This is a flowchart of a side-by-side image detection method according to an embodiment of the present invention;
[0009] Figure 3 This is a flowchart of a side-by-side image detection method according to an embodiment of the present invention;
[0010] Figure 4A This is a schematic diagram illustrating the detection of a second image using an object detection model according to an embodiment of the present invention;
[0011] Figure 4B and Figure 4C This is a schematic diagram illustrating the detection of a second image using a semantic segmentation model according to an embodiment of the present invention;
[0012] Figure 5 This is a flowchart of a side-by-side image detection method according to an embodiment of the present invention;
[0013] Figure 6 This is a schematic diagram of acquiring processed training images according to an embodiment of the present invention.
[0014] Explanation of reference numerals in the attached figures
[0015] 10: Electronic devices;
[0016] 120: Storage device;
[0017] 130: Processor;
[0018] 131: Image acquisition component;
[0019] 132: Side-by-side image detection component;
[0020] 133: Model training component;
[0021] Img1_1, Img1_2, Img1_3: First image;
[0022] P1: Desktop content;
[0023] P2: Browser interface;
[0024] P3, P5: Side-by-side images;
[0025] P3_1, P4_1, P5_1: Left eye image;
[0026] P3_2, P4_2, P5_2: Right eye images;
[0027] Obj1: Detected object;
[0028] R1: Rectangular image block;
[0029] Img6: Original training image;
[0030] Img6_1, Img6_2, Img6_3: Processed training images;
[0031] S210~S220, S310~S324, S510~S550: Steps. Detailed Implementation
[0032] Reference will now be made in detail to exemplary embodiments of the invention, examples of which are illustrated in the accompanying drawings. Wherever possible, the same component reference numerals are used in the drawings and description to denote the same or similar parts.
[0033] Figure 1A and Figure 1B This is a schematic diagram of an electronic device according to an embodiment of the present invention. Please refer to... Figure 1A and Figure 1B The electronic device 10 may include a storage device 120 and a processor 130. The processor 130 is coupled to the storage device 120. In one embodiment, the electronic device 10 may form a 3D display system with a 3D display 20. The 3D display 20 is a stereoscopic display device that can be connected to the electronic device 10 via a wired or wireless data transmission interface. The 3D display 20 may be, for example, a naked-eye 3D display or a glasses-type 3D display. Alternatively, the 3D display 20 may be a head-mounted display device or a computer screen, desktop screen, or television that provides 3D image display capabilities. The 3D display system may be a single integrated system or a discrete system. Specifically, the 3D display 20, storage device 120, and processor 130 in the 3D display system may be implemented as an all-in-one (AIO) electronic device, such as a head-mounted display device, a laptop computer, or a tablet computer, etc., and the processor 130 of the electronic device 10 may be connected to the 3D display 20 via a system bus. Alternatively, the 3D display 20 can be connected to the processor 130 of a computer system via a wired or wireless transmission interface, such as a head-mounted display, desktop screen, or electronic billboard connected to a computer system. The 3D display 20 can receive image data provided by the processor 130 to play 3D images for the user to view based on the image data provided by the processor 130 of the electronic device 10.
[0034] Storage device 120 is used to store images, data, and program code (such as operating system, application program, driver) accessible to processor 130. It can be, for example, any type of fixed or removable random access memory (RAM), read-only memory (ROM), flash memory, hard disk, or a combination thereof.
[0035] Processor 130 is coupled to storage device 120, and is, for example, a central processing unit (CPU), application processor (AP), or other programmable general-purpose or special-purpose microprocessor, digital signal processor (DSP), image signal processor (ISP), graphics processing unit (GPU), or other similar device, integrated circuit, or combination thereof. In one embodiment, processor 130 may access and execute program code and software modules recorded in storage device 120 to implement the side-by-side image detection method in this embodiment of the invention.
[0036] Please refer to Figure 1BIn one embodiment, the processor 130 may include an image acquisition unit 131, a side-by-side image detection unit 132, and a model training unit 133. In one embodiment, the image acquisition unit 131, the side-by-side image detection unit 132, and the model training unit 133 may be implemented in hardware. For example, they may be constructed as components within the processor 130 using dedicated hardware implementation methods such as application-specific integrated circuits (ASICs), programmable logic arrays, and other hardware components. More specifically, the blocks of the image acquisition unit 131, the side-by-side image detection unit 132, and the model training unit 133 may be implemented as logic circuits on an integrated circuit. The related functions of the image acquisition unit 131, the side-by-side image detection unit 132, and the model training unit 133 may be implemented in hardware using hardware description languages (such as Verilog HDL or VHDL) or other suitable programming languages. For example, the functions of the image acquisition component 131, the side-by-side image detection component 132, and the model training component 133 can be implemented in one or more controllers, microcontrollers, microprocessors, application-specific integrated circuits (ASICs), digital signal processors (DSPs), and / or field-programmable gate arrays (FPGAs). That is, the image acquisition component 131, the side-by-side image detection component 132, and the model training component 133 can be multiple hardware components installed within the electronic device 10. In one embodiment, the image acquisition component 131, the side-by-side image detection component 132, and the model training component 133 can be multiple special hardware circuits, and the processor 130 is a special processor architecture integrating multiple special hardware circuits. The image acquisition component 131, the side-by-side image detection component 132, and the model training component 133 have specific connection relationships, and they can be interconnected through their respective data output interfaces and data input interfaces.
[0037] In one embodiment, an image acquisition unit 131 is connected to a storage device 120 and configured to acquire a first image with a first image size from the storage device 120 and output the first image. The image acquisition unit 131 outputs the first image to a side-by-side image detection unit 132. The side-by-side image detection unit 132 is connected to the image acquisition unit 131 and the storage device 120, and receives the first image from the image acquisition unit 131 and a convolutional neural network model acquired from the storage device 120. The side-by-side image detection unit 132 is configured to use the convolutional neural network model to detect a second image conforming to a side-by-side image format in the first image and output the second image, wherein the second image has a second image size. The side-by-side image detection unit 132 is connected to a 3D display 20. The side-by-side image detection unit 132 can directly output the second image conforming to the side-by-side image format to the 3D display 20, or perform other 3D image processing (e.g., image weaving processing required for a naked-eye 3D display) on the second image conforming to the side-by-side image format before outputting the processing result to the 3D display 20. Accordingly, the 3D display 20 can perform 3D display based on the image data provided by the side-by-side image detection unit 132. In this way, by setting multiple special hardware components within the electronic device 10, the electronic device 10 can detect a second image conforming to the side-by-side image format from the first image, thereby improving the user experience and application scope of 3D display technology.
[0038] In some embodiments, the model training unit 133 obtains original training images conforming to a side-by-side image format from the storage device 120, and performs image cropping processing on the original training images to obtain at least one processed training image. The model training unit 133 trains a convolutional neural network model based on the original training images and at least one processed training image, and outputs the convolutional neural network model to the storage device 120.
[0039] Figure 2 This is a flowchart of a side-by-side image detection method according to an embodiment of the present invention. Please refer to... Figure 2 The method of this embodiment is applicable to the electronic device 10 in the above embodiment. The following describes the detailed steps of this embodiment in conjunction with the various components in the electronic device 10.
[0040] In step S210, the processor 130 acquires a first image having a first image size. In one embodiment, this first image may be an image acquired by performing a screen capture function on the screen displayed on the display. The first image may, for example, be image content provided by an application operating in full-screen mode, but the present invention is not limited thereto. The first image may or may not include the user interface of the application. For example, the first image may be a photograph played by a photo player in full-screen mode. Alternatively, the first image may also include a browser interface and image content played by the browser. Furthermore, in one embodiment, this first image may also be a single frame image in an image stream.
[0041] In step S220, processor 130 uses a convolutional neural network (CNN) model to detect a second image in the first image that conforms to a side-by-side image format, wherein the second image has a second image size. The side-by-side (SBS) image format is a 3D image format. The second image conforming to the side-by-side image format includes a left-eye image and a right-eye image arranged horizontally. Here, the trained convolutional neural network model is a deep learning model pre-constructed based on a training dataset for machine learning, and it can be stored in storage device 120. In other words, the model parameters of the trained convolutional neural network model (e.g., the number of neural network layers and the weights of each neural network layer, etc.) have been determined by pre-training and stored in storage device 120.
[0042] In some embodiments, the size of the first image is the same as the size of the second image. In other words, the processor 130 can use a trained convolutional neural network model to determine whether the first image is a second image conforming to a side-by-side image format. Alternatively, in some embodiments, the size of the first image is larger than the size of the second image. In other words, the processor 130 can use a trained convolutional neural network model to determine whether the first image includes a second image conforming to a side-by-side image format, where the second image is a portion of the first image. Therefore, the processor 130 can detect a second image conforming to a side-by-side image format within the first image using a convolutional neural network.
[0043] Therefore, in some embodiments, in response to the processor 130 acquiring a second image conforming to a side-by-side image format from the first image, the processor 130 can control the 3D display to automatically display the second image according to the corresponding playback mode, so as to correctly play the 3D picture that the user wants to watch. Alternatively, in response to the processor 130 acquiring a second image conforming to a side-by-side image format from the first image, the processor 130 can first convert the second image conforming to the side-by-side image format into a 3D format image conforming to another 3D image format, and then control the 3D display to start the 3D display function to play the 3D format image conforming to the other 3D image format. Or, in response to the processor 130 determining that the first image does not include a second image conforming to a side-by-side image format, the processor 130 can generate a specific image conforming to the side-by-side image format based on the image content of the first image, so that the 3D display can play the 3D picture according to the corresponding playback mode.
[0044] Furthermore, in some embodiments, the processor 130 may first determine the contextual attributes of the first image and use a convolutional neural network model corresponding to the contextual attributes to detect image content conforming to the side-by-side image format. The aforementioned contextual attributes may include, for example, cartoon animation attributes, game screen attributes, and real-world scene attributes. In other words, the storage device 120 may record multiple convolutional neural network models corresponding to multiple contextual attributes, each trained on a different training dataset. In some embodiments, the processor 130 may first determine the contextual attributes of the first image and then select one of the multiple convolutional neural network models based on the contextual attributes of the first image for subsequent detection actions. This improves the detection accuracy of side-by-side images. In other words, the processor 130 can train multiple convolutional neural network models for image content with different contextual attributes to further optimize detection accuracy, which is difficult to achieve with traditional image processing techniques.
[0045] Figure 3 This is a flowchart of a side-by-side image detection method according to an embodiment of the present invention. Please refer to... Figure 3 The method described in this embodiment is applicable to the electronic device 10 in the above embodiments. The detailed steps of this embodiment will be explained below with reference to the various components in the electronic device 10. It should be noted that in some embodiments, the side-by-side image detection component 132 may include multiple hardware components to achieve... Figure 3 The steps shown.
[0046] In step S310, processor 130 acquires a first image having a first image size. In step S320, processor 130 uses a convolutional neural network model to detect a second image in the first image that conforms to a side-by-side image format, wherein the second image has a second image size. In this embodiment, step S320 can be implemented as steps S321 to S324.
[0047] In step S321, processor 130 inputs the first image to a convolutional neural network model and obtains a confidence parameter based on the model output data of the convolutional neural network model. The convolutional neural network model includes multiple convolutional layers that perform convolution operations, and may be, for example, an object detection model or a semantic segmentation model. Here, processor 130 can use the convolutional neural network model to detect rectangular image blocks in the first image that may conform to the side-by-side image format. Based on the model output data associated with this rectangular image block, processor 130 can obtain a confidence parameter corresponding to this rectangular image block. In some embodiments, side-by-side image detection unit 132 inputs the first image to the convolutional neural network model and obtains a confidence parameter based on the model output data of the convolutional neural network model. For example, side-by-side image detection unit 132 may include a first unit that receives the first image from image acquisition unit 131 and acquires the convolutional neural network model from storage device 120. The first component can input the first image into the convolutional neural network model, and obtain the confidence parameter based on the model output data of the convolutional neural network model, so as to output the confidence parameter and the model output data of the convolutional neural network model.
[0048] In some embodiments, when the convolutional neural network model is an object detection model, the rectangular image block represents the detected object. Correspondingly, the confidence parameter can be the object classification probability of the detected object, or other parameters generated based on the object classification probability. On the other hand, when the convolutional neural network model is a semantic segmentation model, the rectangular image block represents the image block containing multiple pixels that the semantic segmentation model classifies as belonging to the side-by-side image category. Correspondingly, the confidence parameter can be the pixel density of the multiple pixels in the rectangular image block that are classified as belonging to the side-by-side image category.
[0049] In step S322, the processor 130 determines whether the confidence parameter is greater than a threshold value, which can be set according to actual needs. Specifically, the convolutional neural network model can be used to detect rectangular image blocks in the first image that may conform to the side-by-side image format. When the confidence parameter corresponding to the rectangular image block is greater than the threshold value, the processor 130 can confirm that this rectangular image block is a second image conforming to the side-by-side image format. Conversely, when the confidence parameter corresponding to the rectangular image block is not greater than the threshold value, the processor 130 can confirm that this rectangular image block is not a second image conforming to the side-by-side image format.
[0050] In some embodiments, the side-by-side image detection component 132 may include a second component connected to the first component to receive a confidence parameter in order to determine whether the confidence parameter is greater than a threshold value.
[0051] If step S322 is affirmative, in step S323, reflecting that the confidence parameter is greater than the threshold value, the processor 130 obtains the second image conforming to the side-by-side image format based on the model output data of the convolutional neural network model. Specifically, after confirming that the rectangular image block detected by the convolutional neural network model is the second image conforming to the side-by-side image format, the processor 130 can obtain the block position of the rectangular image block based on the model output data of the convolutional neural network model, and thus obtain the image position of the second image conforming to the side-by-side image format within the first image based on the block position of the rectangular image block. Conversely, if step S322 is negative, in step S324, reflecting that the confidence parameter is not greater than the threshold value, the processor 130 determines that the first image does not include the second image conforming to the side-by-side image format. Therefore, it can be seen that when the first image includes a local image block conforming to the side-by-side image format and other image content, the processor 130 can still accurately detect the local image block conforming to the side-by-side image format using the convolutional neural network model, which is difficult to achieve with traditional image processing techniques.
[0052] In some embodiments, the side-by-side image detection unit 132, in response to a confidence parameter greater than a threshold value, acquires a second image conforming to the side-by-side image format based on the model output data of the convolutional neural network model. In response to a confidence parameter not greater than the threshold value, the processor 130 determines that the first image does not include the second image conforming to the side-by-side image format. In some embodiments, the side-by-side image detection unit 132 may include a third unit connected to the second unit and receiving the output of the second unit to determine whether the confidence parameter is greater than the threshold value. The third unit is connected to the first unit and receives the model output data of the convolutional neural network model output by the first unit. In response to a confidence parameter greater than the threshold value, the third unit acquires a second image conforming to the side-by-side image format based on the model output data of the convolutional neural network model. In response to a confidence parameter not greater than the threshold value, the third unit determines that the first image does not include the second image conforming to the side-by-side image format.
[0053] In some embodiments, the convolutional neural network model includes an object detection model, such as R-CNN, Fast R-CNN, Faster R-CNN, YOLO, or SSD, etc., for object detection; this invention is not limited thereto. The model output data of the object detection model may include the object category, object location, and object classification probability (also known as classification confidence). Based on this, in some embodiments, the confidence parameter may include the object classification probability of the detected object detected by the convolutional neural network model. Furthermore, in some embodiments, the processor 130 may obtain the image position of the second image in the first image based on the object location of the detected object detected by the convolutional neural network model. In some embodiments, the side-by-side image detection unit 132 may obtain the image position of the second image in the first image based on the object location of the detected object detected by the convolutional neural network model.
[0054] Figure 4A This is a schematic diagram illustrating the detection of a second image using an object detection model according to an embodiment of the present invention. Please refer to... Figure 4A The processor 130 can acquire a first image Img1_1 using screen acquisition technology. The first image Img1_1 includes the operating system desktop content P1, the browser interface P2, and a side-by-side image P3 played by the browser. In this example, the side-by-side image P3 conforms to a side-by-side image format and includes a left-eye image P3_1 and a right-eye image P3_2. The processor 130 can input the first image Img1_1 into a trained object detection model. Thereby, the processor 130 can detect objects Obj1 in the first image Img1_1 that may conform to the side-by-side image format using the object detection model, and generate the object position and object classification probability of the detected object Obj1. Then, the processor 130 can determine whether the object classification probability of the detected object Obj1 is greater than a threshold value. If the object classification probability of detected object Obj1 is greater than a threshold value, processor 130 can obtain the image position of a side-by-side image P3 (i.e., the second image) conforming to the side-by-side image format within the first image Img1_1 based on the object position of detected object Obj1. Therefore, processor 130 can detect the second image conforming to the side-by-side image format and can obtain the second image conforming to the side-by-side image format from the first image Img1_1.
[0055] In some embodiments, the convolutional neural network model includes a semantic segmentation model. The model output data of the object detection model may include the classification result of each pixel in the input image. Based on this, in some embodiments, the confidence parameter may include the pixel density of multiple pixels that are determined by the convolutional neural network model to belong to a first category. Furthermore, in some embodiments, the processor 130 may obtain the image position of the second image in the first image based on the pixel positions of the multiple pixels that are determined by the convolutional neural network model to belong to the first category.
[0056] Figure 4B This is a schematic diagram illustrating the detection of a second image using a semantic segmentation model according to an embodiment of the present invention. Please refer to... Figure 4B The processor 130 can acquire a first image Img1_2 from an image stream. In this example, the first image Img1_2 conforms to a side-by-side image format and includes a left-eye image P4_1 and a right-eye image P4_2. The processor 130 can input the first image Img1_2 into a trained semantic segmentation model. The semantic segmentation model can perform a classification operation on each pixel in the first image Img1_2 to obtain a classification result for each pixel in the first image Img1_2. In one embodiment, each pixel in the first image Img1_2 can be classified by the semantic segmentation model into a first category and a second category. The first category represents that the pixel belongs to an image conforming to the side-by-side image format, while the second category represents that the pixel does not belong to an image conforming to the side-by-side image format. The model output data of the semantic segmentation model is the classification result for each pixel in the first image Img1_2.
[0057] At Figure 4B In this example, processor 130 can then calculate the pixel density of the multiple pixels in the first image Img1_2 that are determined to belong to the first category, to obtain a confidence parameter. Specifically, assuming the first image Img1_2 includes N1 pixels, and the number of pixels in the first image Img1_2 that are determined to belong to the first category by the convolutional neural network model is M1, then processor 130 can calculate the pixel density M1 / N1 to obtain the confidence parameter. Processor 130 can determine that the first image Img1_2 is a second image conforming to the side-by-side image format if the confidence parameter is greater than a threshold value. It is worth mentioning that by comparing the threshold value and the confidence parameter, processor 130 can avoid misclassifying the first image Img1_2, which has high image content repetition, as conforming to the side-by-side image format.
[0058] Figure 4C This is a schematic diagram illustrating the detection of a second image using a semantic segmentation model according to an embodiment of the present invention. Please refer to... Figure 4C The first image Img1_3 includes side-by-side image P5 and other image content. In this example, the first image Img1_3 conforms to a side-by-side image format and includes a left-eye image P5_1 and a right-eye image P5_2. The processor 130 can input the first image Img1_3 into the trained semantic segmentation model. Similar to Figure 4B The output data of the semantic segmentation model is the classification result of each pixel in the first image Img1_3.
[0059] Therefore, the processor 130 can determine the rectangular image block R1 from the first image Img1_3 based on the distribution positions of multiple pixels classified as the first category within the first image Img1_3. In some embodiments, the processor 130 can determine the block position of the rectangular image block R1 based on the pixel positions of the multiple pixels classified as the first category. In some embodiments, the block position of the rectangular image block R1 is determined based on the pixel positions of pixels identified by the semantic segmentation model as belonging to the first category. For example, the processor 130 can determine the rectangular image block R1 based on the maximum and minimum X-coordinates, maximum and minimum Y-coordinates of the multiple pixels classified as the first category in the first image Img1_3. Alternatively, in some embodiments, by searching inwards from the four boundaries of the first image Img1_3, the processor 130 can also determine the four boundaries of the rectangular image block R1 based on the pixel positions of pixels identified as belonging to the first category.
[0060] Next, the processor 130 can calculate the pixel density of the multiple pixels in the rectangular image block R1 that are determined by the semantic segmentation model to belong to the first category, in order to obtain a confidence parameter. Specifically, assuming that the rectangular image block R1 includes N2 pixels, and the number of pixels in the rectangular image block R1 that are determined by the semantic segmentation model to belong to the first category is M2, then the processor 130 can calculate the pixel density M2 / N2 to obtain the confidence parameter. Figure 4C In this example, processor 130 can determine that a rectangular image block R1 in the first image Img1_3 conforms to a side-by-side image format if the confidence parameter is greater than a threshold value; that is, rectangular image block R1 is a second image that conforms to the side-by-side image format and has a second image size. Therefore, processor 130 can obtain the image position of the second image conforming to the side-by-side image format in the first image Img1_3 based on the block position of rectangular image block R1. As mentioned above, the block position of rectangular image block R1 is determined based on the pixel positions of some pixels that are determined by the semantic segmentation model to belong to the first category. It is worth noting that by comparing the threshold value and the confidence parameter, processor 130 can avoid misjudging rectangular image blocks R1 with high image content repetition as conforming to the side-by-side image format.
[0061] In some embodiments, when the convolutional neural network model is a semantic segmentation model, the side-by-side image detection unit 132 can extract rectangular image blocks from the first image based on the model output data of the convolutional neural network model. The side-by-side image detection unit 132 can calculate the pixel density of multiple pixels in the rectangular image block that are determined by the convolutional neural network model to belong to the first category, in order to obtain a confidence parameter. The side-by-side image detection unit 132 can obtain the image position of the second image in the first image based on the block position of the rectangular image block, wherein the block position is determined based on the pixel positions of multiple pixels that are determined by the convolutional neural network model to belong to the first category.
[0062] Figure 5 This is a flowchart of a side-by-side image detection method according to an embodiment of the present invention. Please refer to... Figure 5 The method of this embodiment is applicable to the electronic device 10 in the above embodiment. The following describes the detailed steps of this embodiment in conjunction with the various components in the electronic device 10.
[0063] In step S510, the processor 130 acquires the original training image conforming to the side-by-side image format, that is, the original training image including the left eye image and the right eye image.
[0064] In step S520, the processor 130 performs image cropping on the original training image to obtain at least one processed training image. Here, the processor 130 performs data augmentation on the original training image to obtain multiple processed training images. Data augmentation is a way to increase the training dataset, primarily achieved by modifying the original training images.
[0065] It should be noted that, in order to crop out image content that also conforms to the side-by-side image format, in some embodiments, the processor 130 crops out the central region of the side-by-side image to obtain another side-by-side image. Figure 6 This is a schematic diagram illustrating the acquisition of processed training images according to an embodiment of the present invention. Please refer to... Figure 6 After acquiring the original training image Img6 conforming to the side-by-side image format, the processor 130 can obtain processed training images Img6_1, Img6_2, and Img6_3 through image cropping. Processed training image Img6_1 is the left-eye image of the original training image Img6, processed training image Img6_2 is the right-eye image of the original training image Img6, and processed training image Img6_3 is the middle region image of the original training image Img6. Therefore, it can be seen that both the original training image Img6 and the processed training image Img6_3 are side-by-side images conforming to the side-by-side image format, while processed training images Img6_1 and Img6_2 are not.
[0066] After these processed training images are generated through data augmentation operations, the solution objects in the original training images and the solution objects in at least one of the processed training images are selected and assigned a solution category.
[0067] In step S530, processor 130 trains a convolutional neural network model based on the original training image and at least one processed training image. During the training phase of the convolutional neural network model, processor 130 uses multiple images labeled with correct solutions in the training dataset. Specifically, processor 130 can input the original training image and at least one processed training image into the convolutional neural network model. By comparing the output of the convolutional neural network model with the object information of the solution object, processor 130 will gradually update the weight information of the convolutional neural network model, ultimately establishing a convolutional neural network model that can be used to detect side-by-side images conforming to the side-by-side image format.
[0068] In step S540, processor 130 acquires a first image having a first image size. In step S550, processor 130 uses a convolutional neural network model to detect a second image in the first image that conforms to a side-by-side image format, wherein the second image has a second image size.
[0069] In summary, in this embodiment of the invention, even if the first image includes other image content, a second image conforming to the side-by-side image format can be extracted from the first image using a convolutional neural network model. Furthermore, the convolutional neural network model can be trained on a training dataset with similar image context attributes, thereby achieving higher detection accuracy for specific image context attributes. This detection result can be used in various application scenarios, thereby improving the user experience and application scope of 3D display technology. For example, after accurately acquiring the second image conforming to the side-by-side image format, the 3D display can automatically switch to an appropriate image playback mode, thereby enhancing the user experience.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for detecting side-by-side images, characterized in that, include: Obtain a first image with a first image size; as well as A second image conforming to a side-by-side image format is detected in the first image using a convolutional neural network model, wherein the second image has a second image size. The step of detecting the second image in the first image that conforms to the side-by-side image format using the convolutional neural network model includes: The first image is input into the convolutional neural network model, and a confidence parameter is obtained based on the model output data of the convolutional neural network model; and If the confidence parameter is greater than the threshold value, the second image conforming to the side-by-side image format is obtained based on the model output data of the convolutional neural network model. The convolutional neural network model includes a semantic segmentation model, and the step of obtaining the confidence parameter based on the model output data of the convolutional neural network model includes: A rectangular image block is obtained from the first image based on the model output data of the semantic segmentation model, wherein the model output data of the semantic segmentation model is the classification result of each pixel in the first image, and the pixel positions of multiple pixels classified into the first category are used to obtain the block position of the rectangular image block; and The pixel density of multiple pixels in the rectangular image block that are determined by the convolutional neural network model to belong to the first category is calculated to obtain the confidence parameter.
2. The side-by-side image detection method according to claim 1, characterized in that, The first image size is the same as the second image size.
3. The side-by-side image detection method according to claim 1, characterized in that, The first image size is larger than the second image size.
4. The side-by-side image detection method according to claim 1, characterized in that, The step of detecting the second image in the first image that conforms to the side-by-side image format using the convolutional neural network model further includes: If the confidence parameter is not greater than the threshold value, it is determined that the first image does not include the second image that conforms to the side-by-side image format.
5. The side-by-side image detection method according to claim 1, characterized in that, The convolutional neural network model includes an object detection model, and the confidence parameter includes the object classification probability of the detected objects by the convolutional neural network model.
6. The side-by-side image detection method according to claim 5, characterized in that, The step of obtaining the second image conforming to the side-by-side image format based on the model output data of the convolutional neural network model includes: The image position of the second image in the first image is obtained based on the object position detected by the convolutional neural network model.
7. The side-by-side image detection method according to claim 1, characterized in that, The step of obtaining the second image conforming to the side-by-side image format based on the model output data of the convolutional neural network model includes: The image position of the second image in the first image is obtained based on the block position of the rectangular image block, wherein the block position is determined based on the pixel position of the portion of pixels that are determined by the convolutional neural network model to belong to the first category.
8. The side-by-side image detection method according to claim 1, characterized in that, The method further includes: Obtain the original training images that conform to the side-by-side image format; The original training images are cropped to obtain at least one processed training image; and The convolutional neural network model is trained based on the original training image and the at least one processed training image.
9. An electronic device, characterized in that, include: Storage device, containing multiple modules; as well as The processor, connected to the storage device, is configured to: Obtain a first image with a first image size; as well as A second image conforming to a side-by-side image format is detected in the first image using a convolutional neural network model, wherein the second image has a second image size. The processor is further configured to: The first image is input into the convolutional neural network model, and the confidence parameter is obtained based on the model output data of the convolutional neural network model. as well as If the confidence parameter is greater than the threshold value, the second image conforming to the side-by-side image format is obtained based on the model output data of the convolutional neural network model. The convolutional neural network model includes a semantic segmentation model, and the processor is configured to: A rectangular image block is obtained from the first image based on the model output data of the semantic segmentation model, wherein the model output data of the semantic segmentation model is the classification result of each pixel in the first image, and the pixel positions of multiple pixels classified as the first category are used to obtain the block position of the rectangular image block. as well as The pixel density of multiple pixels in the rectangular image block that are determined by the convolutional neural network model to belong to the first category is calculated to obtain the confidence parameter.
10. The electronic device according to claim 9, characterized in that, The first image size is the same as the second image size.
11. The electronic device according to claim 9, characterized in that, The first image size is larger than the second image size.
12. The electronic device according to claim 9, characterized in that, The processor is configured to: determine that the first image does not include the second image that conforms to the side-by-side image format if the confidence parameter is not greater than the threshold value.
13. The electronic device according to claim 9, characterized in that, The convolutional neural network model includes an object detection model, and the confidence parameter includes the object classification probability of the detected objects by the convolutional neural network model.
14. The electronic device according to claim 13, characterized in that, The processor is configured to: obtain the image position of the second image in the first image based on the object position of the detected object detected by the convolutional neural network model.
15. The electronic device according to claim 9, characterized in that, The processor is configured to: obtain the image position of the second image in the first image based on the block position of the rectangular image block, wherein the block position is determined based on the pixel positions of a plurality of pixels that are determined by the convolutional neural network model to belong to the first category.
16. The electronic device according to claim 9, characterized in that, The processor is configured to: Obtain the original training images that conform to the side-by-side image format; The original training image is cropped to obtain at least one processed training image; as well as The convolutional neural network model is trained based on the original training image and the at least one processed training image.