Image processing method and intelligent terminal
Through the smart terminal to acquire and process images in real time, binocular images are generated, combined with gesture control and element replacement functions, the problem of high-cost VR equipment is solved, and a low-cost real-time three-dimensional interactive experience is achieved.
Patent Information
- Application Number
- CN202211375460.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-04
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-11-04
AI Technical Summary
Existing MR equipment and VR equipment are costly, and simple VR equipment requires a dedicated chip source to achieve interactive experience, and lacks real-time interaction functions.
The intelligent terminal collects environmental images in real time, generates binocular images and loads displays, uses gesture control, element replacement and device control function modules to achieve real-time interaction, and combines cloud or local image processing to generate three-dimensional three-dimensional effects.
It realizes a low-cost real-time three-dimensional interactive experience, avoids the use of high-cost VR devices, simplifies the interaction process, and saves user costs.
Smart Images

Figure CN115661414B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and in particular relates to an image processing method and an intelligent terminal. Background Art
[0002] At present, with the development of science and technology and the improvement of people's living standards, more and more users are willing to experience new technological products. Among them, MR (Mixed Reality) devices and VR (Virtual Reality) devices can achieve a three-dimensional stereoscopic effect on the image currently seen by the user through their image acquisition and display processing technology, thereby realizing the user's mixed reality or virtual reality experience. The current development of MR devices and VR devices is in a good situation, and MR devices and VR devices on the market are also emerging in an endless stream.
[0003] However, due to the hardware costs of MR and VR devices, the existing MR and VR devices are relatively expensive. The prices of MR or VR devices that carry cameras for interaction are even higher, making them difficult to popularize. Although the simple VR device formed by combining a smart terminal with a VR box has a lower cost, it requires downloading a video source with a three-dimensional effect for interactive experience. However, there are currently few VR videos available for playback on the market, and the production cost of specially produced VR videos is also high. At the same time, it does not have the function of real-time interaction and can only realize video playback, resulting in a simple interactive function. Summary of the Invention
[0004] The purpose of the embodiments of the present invention is to provide an image processing method, aiming to solve the problem of high cost of existing MR equipment and VR equipment.
[0005] The embodiment of the present invention is implemented as follows: an image processing method is applied to a smart terminal, the method comprising:
[0006] Real-time collection of image information in the current environment;
[0007] Performing image processing on the currently acquired image information to generate a binocular image, wherein the binocular image is a binocular view image in which the left view and the right view are located on the left and right sides respectively;
[0008] Load and display the currently generated binocular image.
[0009] Furthermore, before the step of real-time acquisition of image information in the current environment, the method further includes:
[0010] The states and parameters set by the user for various functional modules are obtained, wherein the functional modules include a gesture control functional module, an element replacement functional module, and a device control functional module.
[0011] Furthermore, when it is obtained that the user sets the gesture control function module to be enabled, the step of performing image processing based on the currently collected image information to generate a binocular image includes:
[0012] Identify whether there is preset gesture information in the collected image information;
[0013] If yes, a preset number of function module setting buttons are loaded into the image information, and image processing is performed to generate a binocular image;
[0014] If not, the image processing is directly performed based on the image information to generate a binocular image.
[0015] Furthermore, when it is obtained that the user sets the element replacement function module to be enabled, the step of performing image processing based on the currently collected image information to generate a binocular image includes:
[0016] Identify the characteristic information of each person contained in the collected image information;
[0017] According to the parameters set by the user for the element replacement function module, the human features are added or synthesized and replaced with virtual elements, and image processing is performed to generate a binocular image.
[0018] Furthermore, after the step of adding or synthesizing the character features into virtual elements according to the parameters set by the user for the element replacement function module, the method further includes:
[0019] Identifying whether there is switching gesture information in the collected image information;
[0020] If so, the virtual element is switched accordingly according to the switching gesture.
[0021] Furthermore, when it is obtained that the user has set the device control function module to be enabled, the method further includes:
[0022] Identifying whether there is control gesture information in the collected image information;
[0023] If so, the operation of the connected device is controlled accordingly according to the parameters set by the user for the device control function module.
[0024] Furthermore, after the step of loading a preset number of function module setting buttons into the image information, the method further includes:
[0025] When a user clicks a button on any function module, the state and parameters of the clicked function module are loaded;
[0026] Modify the functional modules accordingly according to the status and parameters set by the user.
[0027] Furthermore, the step of performing image processing to generate a binocular image based on the currently collected image information includes:
[0028] Upload the currently collected image information to the cloud server in real time and receive the binocular image returned by the cloud server after real-time image processing based on the image information; or
[0029] The currently acquired image information is processed locally to generate a binocular image.
[0030] Furthermore, the step of performing image processing to generate a binocular image based on the currently collected image information includes:
[0031] Performing image processing on two sets of images captured by a binocular camera at different viewing angles to generate a binocular image; or
[0032] The single set of images captured by the monocular camera is processed to generate a virtual right image, and the single set of images is compared with the virtual right image. Figure 1 Perform image processing to generate a binocular image; or
[0033] The single set of images captured by the monocular camera is processed to generate a binocular image.
[0034] Another embodiment of the present invention aims to provide an intelligent terminal, comprising a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned image processing method when executing the program.
[0035] The image processing method provided by the embodiment of the present invention is applied to a smart terminal. After the smart terminal is loaded into a VR box, the smart terminal collects image information of the current environment in real time, performs image processing on the currently collected image information to generate a binocular image, and then loads and displays the currently generated binocular image, that is, the binocular image is displayed on the screen of the smart terminal. At this time, when the user uses the VR box to view the binocular view, a three-dimensional stereo image is synthesized in the human brain, thereby achieving a three-dimensional stereo effect, thereby achieving a real-time interactive experience for the user, thereby avoiding the problem that the simple VR box composed of an existing simple smart terminal and a VR box requires a dedicated film source for the interactive experience. At the same time, the image collected in real time by the smart terminal can be processed to produce a three-dimensional stereo effect, making the interaction easier. At the same time, since the user's own smart terminal is used as a carrier for the interactive experience, the user's usage cost can be effectively saved, solving the problem of high cost of existing MR equipment and VR equipment. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a flowchart of an image processing method provided by an embodiment of the present invention;
[0037] Figure 2 is another flow chart of an image processing method provided by an embodiment of the present invention;
[0038] Figure 3 It is a structural diagram of the intelligent terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0040] In the present invention, unless otherwise expressly specified or limited, the terms "installed," "connected," "connected," "fixed," and the like should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; they may refer to mechanical connections or electrical connections; they may refer to direct connections or indirect connections through an intermediate medium; and they may refer to internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on the specific circumstances. The term "and / or" as used herein includes any and all combinations of one or more of the relevant listed items.
[0041] Example 1
[0042] See also Figure 1 , is a flow chart of an image processing method provided by a first embodiment of the present invention. For ease of description, only the portion related to the embodiment of the present invention is shown. The method includes:
[0043] Step S10, collecting image information of the current environment in real time;
[0044] In one embodiment of the present invention, the image processing method is applied to a smart terminal, specifically a smart phone, wherein the smart terminal needs to be loaded into a VR (Virtual Reality) box when in use to realize the user's virtual reality or mixed reality interactive experience, wherein the VR box is a device that can load and unload the smart terminal and can adjust the pupil distance and focal length.
[0045] The smart terminal is provided with at least one camera for capturing images and videos. Specifically, the smart terminal may be provided with a monocular camera, a binocular camera, or a multi-camera camera, which is determined according to the specific configuration of the smart terminal and is not specifically limited herein.
[0046] Furthermore, after the user uses the smart terminal and loads it into the VR box, the smart terminal collects image information of the current environment in real time, that is, the smart terminal uses the camera it is equipped with to shoot real-time video, and each frame of the video information shot by the smart terminal is the image collected by the smart terminal every hour.
[0047] Among them, it should be pointed out that, when a monocular camera is provided in the smart terminal, the image information collected by the smart terminal in real time includes a single set of images collected by the monocular camera; when a binocular camera is provided in the smart terminal, the image information collected by the smart terminal in real time includes two sets of images collected by the binocular camera at two different perspectives. At this time, the two sets of images have a specific angle difference of two, which is similar to the two eyes of a person; correspondingly, when a multi-camera is provided in the smart terminal, the image information collected by the smart terminal in real time includes multiple sets of images collected by the multi-camera at multiple different perspectives.
[0048] Step S20, performing image processing based on the currently collected image information to generate a binocular image;
[0049] Among them, in one embodiment of the present invention, after the smart terminal collects the current image information, the image information is processed accordingly in real time to generate a binocular image, wherein the binocular image is a binocular view image with the left view and the right view located on the left and right sides respectively, that is, the binocular image is a view image in which the left view and the right view are spliced together into a whole. When the user watches the binocular view through a VR box, an interactive experience with a three-dimensional stereo effect can be generated.
[0050] In one embodiment of the present invention, performing image processing based on the currently collected image information to generate a binocular image includes the following three implementation methods:
[0051] 1. Processing two sets of images from different perspectives captured by the binocular camera together to generate a binocular image;
[0052] 2. Process the single set of images captured by the monocular camera to generate a virtual right image, and compare the single set of images with the virtual right image. Figure 1 Perform image processing to generate binocular images;
[0053] 3. Perform image processing on the single set of images captured by the monocular camera to generate a binocular image.
[0054] Specifically, in method 1, when the smart terminal uses a binocular camera, since its binocular camera can capture two sets of images, namely the left view and the right view required in the binocular image, it should be noted that the types of the two sets of images captured by the binocular camera are the same, only the perspectives are different. For example, the binocular camera is two conventional cameras of similar types, wide-angle cameras, or telephoto cameras. At this time, the two sets of images captured by the binocular camera can be directly spliced into a binocular image, that is, the images captured under the perspectives of different cameras are directly loaded on the left and right sides. At this time, due to the parallax between the left and right views in the binocular image, when the user uses a VR box to watch the binocular view, a three-dimensional stereo image will be synthesized in the human brain, thereby achieving a three-dimensional stereo effect. In accordance with the above, when the smart terminal uses a multi-camera, it can also extract the two sets of images required from the multiple sets of images captured by the multi-camera and splice them into a binocular image.
[0055] In the second method, when the smart terminal uses a monocular camera, its monocular camera only captures a single set of images. In order to achieve a better three-dimensional stereo effect, the single set of images captured by the monocular camera is processed to generate a virtual right image, wherein the virtual right image is a virtual image under the virtual perspective of the 3D right image after depth conversion processing based on the foreground image captured by the monocular camera. Specifically, in an embodiment of the present invention, the virtual right image is generated by performing depth processing on the foreground image to obtain a depth map (also called an optical flow map) with depth information, and then processing the depth map to improve the contrast of the main part. It dynamically converts the depth map into a mask, and then uses the mask to deduct the main body, and rotates, stretches, and shifts all pixels in the deducted main body based on the depth and fills them back to the original image after rotating, stretching, and shifting to the left, and then performing image repair on the hole area generated by the left shift (for example, the mean filling method), and finally generating a virtual right image that smoothes the entire image and erases small "cracks". It should be pointed out that the image processing method of the virtual right image can also be implemented by any method in the prior art, which is not specifically limited here.
[0056] Furthermore, after generating the virtual right image, the single group of images is compared with the virtual right image. Figure 1 Image processing is then performed to generate a binocular image. This is done using the same method described above, loading a single image set and a virtual right image on each side to create a binocular image. This stitching of the single image set and the newly generated virtual right image together simulates the 3D perspective normally seen by the human eye. Accordingly, when the user views this binocular view through a VR headset, a 3D image is synthesized in the user's brain, achieving a 3D effect.
[0057] In the third method, when the smart terminal uses a monocular camera, its monocular camera collects a single set of images. At this time, the image processing is directly performed using the above-mentioned method, so that the monocular image is copied and spliced into left and right layers, thereby presenting a completely identical image of the left view and the right view in the binocular image. It should be pointed out that the research and investigation found that the human eye judges three-dimensionality mainly through the light and shadow effects. The left view and the right view in the binocular image are completely identical. Figure 1 Under certain circumstances, users can also synthesize three-dimensional stereoscopic images in their brains, but the three-dimensional effect and stereoscopic feeling are not as obvious as those in the above-mentioned methods 1 and 2.
[0058] Therefore, it specifically selects any one of the above three methods to perform image processing to generate a binocular image according to the actual hardware parameter configuration of the smart terminal or the options set by the user. No specific limitation is made here. For example, when the smart terminal adopts a monocular camera and the processing performance is weak or the user's stereoscopic sense requirement is low, it can adopt the above method three to perform image processing to generate a binocular image; when the smart terminal adopts a monocular camera and the processing performance is strong or the user's stereoscopic sense requirement is high, it can adopt the above method two to perform image processing to generate a binocular image; and when the smart terminal adopts a binocular camera or a multi-camera, it can adopt the above method one to perform image processing to generate a binocular image. Of course, according to user needs, the above method two can also be used to perform image processing on a specific set of two or more sets of images captured by the binocular camera or the multi-camera to generate a binocular image to enhance the user's stereoscopic sense when watching.
[0059] Furthermore, it should be pointed out that in a binocular camera or a multi-camera, each camera is a regular camera, a wide-angle camera, or a telephoto camera, rather than two similar cameras with different viewing angles. At this time, the images taken by each camera are inconsistent, so the above-mentioned method 1 cannot be used for image processing. At this time, method 2 needs to be used for image processing to generate a binocular image.
[0060] Furthermore, in one embodiment of the present invention, performing image processing based on the currently collected image information to generate a binocular image also includes the following two implementation methods:
[0061] 1. Upload the currently collected image information to the cloud server in real time and receive the binocular image returned by the cloud server after real-time image processing based on the image information;
[0062] 2. Perform local image calculation processing on the currently collected image information to generate a binocular image.
[0063] Specifically, in method one, the smart terminal can upload the currently collected image information to the cloud server in real time, and use the cloud server to process the image. At this time, when the smart terminal is configured with a 5G network, the network transmission delay is only a few milliseconds. Therefore, when the smart terminal is insufficiently configured and has small computing power, or needs to save power, or the computing power required for image processing is large, the image information can be quickly processed through the large computing power of the cloud server, and a binocular image is generated and returned to the smart terminal after the processing is completed. The smart terminal can directly receive the binocular image returned by the cloud server, which can alleviate the computing power demand of the smart terminal and correspondingly save the energy consumption of the smart terminal.
[0064] In the second method, when the intelligent terminal is equipped with sufficient computing power, or when power consumption is not a concern or the computing power required for image processing is small, the local computing power can be used to directly perform image calculations and generate binocular images.
[0065] Furthermore, in a preferred embodiment of the present invention, it can also be a combination of the above two types of implementation. For example, when the binocular camera collects binocular images, the local computing power of the smart terminal is directly used to perform image calculation and processing to generate a binocular image, and when the monocular camera collects monocular images, the image information is uploaded to the cloud server in real time through the smart terminal for calculation and processing to obtain a binocular image. Alternatively, the cloud server can perform image processing on the monocular image to generate a virtual right image and return it to the smart terminal, and then the smart terminal uses local computing power to process the monocular image and the virtual right image to generate a binocular image. The above implementation methods are specifically combined and set according to actual usage needs, and no specific limitation is made here.
[0066] Furthermore, in one embodiment of the present invention, the image processing based on the currently collected image information to generate a binocular image is specifically implemented by the following steps:
[0067] Convert the two sets of images into arrays respectively; broadcast the arrays into matrices and merge the arrays; and convert the merged arrays into a binocular image.
[0068] Specifically, referring to the above description, the two sets of images can be two sets of images captured by a binocular camera from different perspectives, or a single set of images captured by a monocular camera and a virtual right image generated after image processing, or a single set of images captured by a monocular camera and a copy thereof. By first converting the two sets of images into arrays, and then converting the images into a multidimensional array, the problem of stitching the two sets of images can be converted into the merging of two multidimensional arrays. The merged multidimensional array is then converted into a binocular image, thereby completing the stitching of the two sets of images into a binocular image.
[0069] Step S30, loading and displaying the currently generated binocular image;
[0070] In one embodiment of the present invention, after the smart terminal processes the currently captured image information to generate a binocular image, the smart terminal loads and displays the currently generated binocular image, that is, displays the binocular image on the screen of the smart terminal. At this time, when the user uses a VR box to view the binocular view, a three-dimensional stereo image is synthesized in the human brain, thereby achieving a three-dimensional stereo effect, thereby achieving a real-time interactive experience for the user, that is, the image viewed by the user in real time is generated with a three-dimensional stereo effect. In the existing technology, the interactive experience can only be achieved by downloading a video source with a three-dimensional effect through the smart terminal. However, there are currently few videos available for playback on the market, and the cost of video production is high. At the same time, the real-time interactive function is not available.
[0071] At the same time, users can use their own smart terminals in combination with VR boxes to use smart terminals as a carrier of interactive experience, avoiding the usage cost problem caused by the high price of VR devices when users need to use VR devices for interaction, thereby effectively saving user costs.
[0072] In this embodiment, the image processing method is applied to a smart terminal. After the smart terminal is loaded into a VR box, the smart terminal collects image information of the current environment in real time, processes the currently collected image information to generate a binocular image, and then loads and displays the currently generated binocular image, that is, the binocular image is displayed on the screen of the smart terminal. At this time, when the user uses the VR box to watch the binocular view, a three-dimensional stereo image will be synthesized in the human brain to achieve a three-dimensional stereo effect, thereby realizing a real-time interactive experience for the user, avoiding the problem that the simple VR box composed of the existing simple smart terminal and the VR box requires a dedicated film source for the interactive experience. At the same time, the picture collected in real time by the smart terminal can be processed to produce a three-dimensional stereo effect, making the interaction easier. At the same time, since the user's own smart terminal is used as a carrier of the interactive experience, the user's usage cost can be effectively saved, solving the problem of high cost of existing MR equipment and VR equipment.
[0073] Example 2
[0074] See also Figure 2 , is a flow chart of an image processing method provided by the second embodiment of the present invention. For ease of description, only the portion related to the embodiment of the present invention is shown. The method of the second embodiment is substantially the same as that of the first embodiment. For the sake of brevity, reference may be made to the corresponding content of the first embodiment for matters not mentioned in this embodiment. Specifically, the method includes:
[0075] Step S11 , obtaining the status and parameters set by the user for various functional modules.
[0076] In the embodiment of the present invention, before the user loads the smart terminal into the VR box, they can also set various functional modules, so that after the smart terminal is loaded into the VR box, it will perform corresponding operations according to the user's settings for the various functional modules. Of course, in the current use environment of the smart terminal, the user can also directly load it into the VR box without setting up the various functional modules. In this case, the status and parameters of the various functional modules obtained by the smart terminal are the initial preset status and parameters, or the status and parameters set by the user last time.
[0077] Among them, the functional modules include but are not limited to gesture control functional module, element replacement functional module, device control functional module, screenshot sharing functional module, and information search functional module. It can be understood that according to subsequent functional development, it can also add various other functional modules through module encapsulation to achieve various functional configurations, which are not specifically limited here.
[0078] Specifically, its gesture control function module is mainly used for users to control the above-mentioned various function modules through gesture control, which is similar to the combination of mouse and keyboard in computer equipment; the element replacement function module is mainly used to add or synthesize various character features contained in the image information collected by the camera in the smart terminal and replace them with virtual elements; the device control function module is mainly used to control the connected devices through gesture control; the screenshot sharing function module is mainly used to take screenshots and share the image information collected by the camera in the smart terminal through gesture control; and the information search function module is mainly used to search for information through gesture control.
[0079] Step S21: collecting image information of the current environment in real time.
[0080] Step S31, performing image processing based on the currently collected image information to generate a binocular image;
[0081] In one embodiment of the present invention, when it is obtained that the user has set the gesture control function module to be enabled, the step of performing image processing based on the currently collected image information to generate a binocular image includes:
[0082] Identify whether there is preset gesture information in the collected image information;
[0083] If yes, a preset number of function module setting buttons are loaded into the image information, and image processing is performed to generate a binocular image;
[0084] If not, the image processing is directly performed based on the image information to generate a binocular image.
[0085] In one embodiment of the present invention, when it is obtained that the user has set the element replacement function module to be enabled, the step of performing image processing based on the currently collected image information to generate a binocular image includes:
[0086] Identify the characteristic information of each person contained in the collected image information;
[0087] According to the parameters set by the user for the element replacement function module, the human features are added or synthesized and replaced with virtual elements, and image processing is performed to generate a binocular image.
[0088] In one embodiment of the present invention, when it is obtained that the user has set the device control function module to be enabled, the method further includes:
[0089] Identifying whether there is control gesture information in the collected image information;
[0090] If so, the operation of the connected device is controlled accordingly according to the parameters set by the user for the device control function module.
[0091] In one embodiment of the present invention, when it is obtained that the user has set the screenshot sharing function module to be enabled, the method further includes:
[0092] Identify whether sharing gesture information exists in the collected image information;
[0093] If yes, take a screenshot of the image in the image information and save it, and prompt the user whether to share it;
[0094] If the user confirms to share, the screenshot will be shared to the social platform bound to the user account.
[0095] Specifically, after the user sets the gesture control function module to be enabled in step S11, the smart terminal recognizes the collected image information in real time. When the image information is identified by image recognition technology as containing preset gesture information (e.g., an OK gesture), it accordingly loads a preset number of function module setting buttons into the image information. Specifically, each function module setting button can be added to the surrounding position of the recognized preset gesture. For example, if a user using a VR box equipped with a smart terminal places their hand in front of the camera and performs a preset action (e.g., an OK gesture), the smart terminal will control the image information to pop up multiple function module setting buttons around the hand. Of course, each function module setting button can also be added to a preset position in the image information. For example, no matter where the user places their hand in the image information, multiple function module setting buttons will pop up at a preset position in the image (e.g., the lower right corner of the image). The user can then click any function module setting button to make it respond accordingly. It should be noted that the preset gesture can be entered by the user or selected by the user from a plurality of preset gestures, and this is not specifically limited here.
[0096] Furthermore, after the step of loading a preset number of function module setting buttons into the image information, the method further includes:
[0097] When a user clicks a button on any function module, the state and parameters of the clicked function module are loaded;
[0098] Modify the functional modules accordingly according to the status and parameters set by the user.
[0099] Specifically, when the user sets the gesture control function module to the enabled state and performs a preset gesture, such as the OK gesture mentioned above, the smart terminal recognizes the preset gesture and loads the various function module setting buttons in the image information accordingly, and at the same time starts to recognize and capture the user's hand accordingly. When the user moves to any function module setting button and clicks with the middle finger (that is, the click gesture information is recognized), the corresponding state and parameters of the clicked function module are loaded in response to the click gesture operation. The specific method of the above-mentioned click gesture recognition is: when the user uses the middle finger to click, there will be a certain gap change between the middle finger and other joints. When the gap change is recognized, it is determined to be the user's click operation, and the corresponding point at the index finger position is clicked for confirmation.
[0100] Specifically, for example, when a user clicks on an element replacement module, the status and parameters of that module are loaded accordingly. The user can then enable or disable the module and set its parameters accordingly. In other words, in addition to setting up each module before loading the smart terminal into the VR box, the user can also set up each module after enabling the gesture control module and performing a preset gesture.
[0101] Furthermore, when the user sets the element replacement function module to be enabled, the smart terminal uses image recognition technology to identify the various character feature information contained in the collected image information, where the character feature information includes face information, head information, hand information, leg information, foot information, etc. At this time, the character features are added or synthesized and replaced with virtual elements according to the parameters set by the user for the element replacement function module. For example, when the parameters set by the user are to replace the faces of all characters in the image with preset virtual faces, the smart terminal synthesizes and replaces all faces in the identified image information with preset virtual faces, and performs image processing to generate a binocular image. The method of generating the binocular image is mainly that each virtual element has a virtual model (i.e., a 3D filter). At this time, the synthesized left image model is loaded in the left view of the binocular image, and the synthesized right image model is loaded in the right view, so that the user can also observe the virtual elements with three-dimensional effects when observing the binocular image.
[0102] Furthermore, after the step of adding or synthesizing the character features into virtual elements according to the parameters set by the user for the element replacement function module, the method further includes:
[0103] Identifying whether there is switching gesture information in the collected image information;
[0104] If so, the virtual element is switched accordingly according to the switching gesture.
[0105] Specifically, when the user uses a switching gesture (for example, the user moves his hand from left to right or from right to left), the smart terminal recognizes the presence of switching gesture information in the image information and switches the virtual elements accordingly according to the switching gesture. For example, when the user moves his hand from left to right, the previous virtual element is switched; and when the user moves his hand from right to left, the next virtual element is switched.
[0106] Furthermore, after the above-mentioned user performs a click operation on the element replacement function module to load the status and parameters of the element replacement function module, the user can further perform a click gesture operation accordingly to adjust the activation status of the virtual element and the various parameters required to be added or synthesized for replacement. In an embodiment of the present invention, the parameters may include, for example, the number of characters that need to replace the virtual elements, the specific virtual elements that need to be replaced, and the specific parameters of the virtual elements.
[0107] Specifically, for example, the character quantity parameter allows users to select either all characters or selected characters. When the user selects all characters, all recognized characters are replaced with the same virtual element. When the user selects selected characters, a selection box is added to each recognized character. When the user selects any of the recognized characters, the virtual element is replaced with the selected character. In other words, the virtual element can be replaced with one or any number of characters selected by the user, while the other characters remain unchanged.
[0108] Specifically, for example, a face, head, hand, leg, foot, etc. can be selected from the specific virtual element parameters. The specific details can be referenced from the filter library in existing image editing programs and are not specifically limited here. At this point, the user can correspondingly add or synthetically replace the face, head, hand, leg, foot, etc. of the person identified in the image information. For example, when the user selects a face or head from the specific virtual element, the face or head is replaced with a pre-selected or currently selected virtual face or headgear. When the user performs a switching gesture, the virtual face or headgear is switched accordingly, thereby achieving the function of virtual face or headgear replacement. Furthermore, when the user selects a hand, leg, foot, etc. from the specific virtual element, virtual wear can be added to the hand, leg, or foot. The virtual wear can be, for example, a virtual watch, virtual bracelet, virtual clothing, virtual shoes, etc. When the user performs a switching gesture, the virtual wear is switched accordingly. In other words, the element replacement function module can implement the addition or synthetic replacement of various virtual elements, such as face or headgear replacement, for the person in the image information.
[0109] Specifically, the specific parameters of the virtual elements are the specific models of the various virtual elements mentioned above, such as specific virtual faces. At this time, the user can switch the virtual faces accordingly through switching gestures. The user can also directly select the specific virtual face to be switched in the specific parameters of the virtual element, so that the character can be directly replaced with the specific virtual face required by the user, without the need for the user to constantly control the switching gestures.
[0110] Therefore, specifically, after the user sets the status and parameters in the element replacement function module, it can replace the same virtual element for all characters in the image information, or replace different virtual elements and their specific parameters for different characters.
[0111] Furthermore, when the user sets the device control function module to be enabled, the smart terminal uses image recognition technology to identify the collected image information. When the image information contains control gesture information (such as the user's snapping gesture) through the image recognition technology, it controls the operation of the connected device according to the parameters set by the user for the device control function module. For example, the smart terminal matches and connects with the external device through the provided WiFi module, Bluetooth module, NFC module, or infrared module. At this time, the smart terminal can directly control the matched external device, such as controlling lamps, televisions, fans, and other devices with infrared control functions through the infrared module, or controlling Bluetooth speakers, televisions, and other devices with specific Bluetooth control functions through the Bluetooth module. In this case, for example, the user sets the parameters of the device control function module to turn the light on and off when the user snaps his fingers. When the user performs the finger snapping operation, the smart terminal performs the light on / off operation accordingly. Correspondingly, after using the VR box, the user can also make a preset gesture first, so that multiple function module setting buttons pop up, and then click the device control function module to pop up the corresponding status and parameters. At this time, the user can set the parameters corresponding to the current control gesture, that is, the operation of the device controlled by each control gesture. It should be pointed out that there can be multiple control gestures, and at this time, any control gesture can control a device corresponding to it; of course, there can also be one control gesture. At this time, when it is necessary to switch the device corresponding to the control of the control gesture, it can be done by changing the parameters of the set device control function module. Of course, it should be pointed out that it can also control the operation of the connected device through voice control.
[0112] Furthermore, when the user sets the screenshot sharing function module to be enabled, the smart terminal identifies the collected image information through image recognition technology, and when the image recognition technology identifies that there is sharing gesture information in the image information (for example, the user makes a heart-shaped gesture with both hands), the smart terminal takes a screenshot and saves the image in the image information. It should be pointed out that the screenshot can be saved when the sharing gesture is recognized. At this time, the screenshot contains the user's sharing gesture. In order to avoid the sharing gesture appearing in the screenshot, the screenshot saving time can also be delayed by a preset time (for example, a delay of two seconds). At this time, the smart terminal starts to pop up a screenshot prompt after recognizing the sharing gesture, and takes a screenshot and saves the image in the currently collected image information after the preset time delay. Furthermore, after completing the screenshot saving operation, the smart terminal can also prompt the user whether to share it. When the user is sure to share it, the screenshot image will be shared to the social platform of the bound user account. It should be pointed out that the user can pre-bind a user account in the smart terminal. At this time, the screenshot image can be directly shared to the social platform of the bound user account; when the user has not bound a user account in the smart terminal and the user is sure to share it, the account login interface will pop up accordingly to enable the user to log in to the account to share the screenshot; of course, the user can also choose to directly share it to the social platform of the bound user account whenever a screenshot operation is performed.
[0113] Furthermore, when the user sets the information search function module to be enabled, multiple groups of information search buttons are loaded accordingly. Specifically, for example, text search can be performed by loading a web page. At this time, the Internet access function can be realized, and a virtual keyboard pops up on the search interface to allow the user to input the content to be searched by hand, and load the search results accordingly when the search is confirmed. Of course, text input can also be performed by recording during the input process. It should be pointed out that all of the above information can be loaded in a preset size and displayed in a preset position in the image information (for example, the lower left corner) to avoid the problem of the display interface being too large and filling the image screen. Of course, it can also perform image search through screenshots or currently collected images. For example, when a user uses a VR box to view a QR code in a scene, it can use the information search function module to scan the QR code to obtain the required information; or when a user uses a VR box to view a person's clothing, it can use the information search function module to search for the clothing information on the e-commerce platform, and at the same time, it can also place an order for shopping; or when a user uses a VR box to view a person's face, it can use the information search function module to search for the person's identity information based on the facial information on a public platform or a social platform, and at the same time, it can also follow the other party's account on a public platform or request to add the other party as a friend on a social platform. It is understandable that the information search method can also be other, which is set according to actual use needs and is not specifically limited here.
[0114] Step S41: loading and displaying the currently generated binocular image.
[0115] Example 3
[0116] A third embodiment of the present invention provides an image processing device. For ease of illustration, only portions related to the present embodiment are shown. The image processing device includes:
[0117] Image acquisition module, used to collect image information of the current environment in real time;
[0118] An image generation module is used to perform image processing based on the currently collected image information to generate a binocular image, wherein the binocular image is a binocular view image with a left view and a right view located on the left and right sides respectively;
[0119] The image loading module is used to load and display the currently generated binocular image.
[0120] Furthermore, in one embodiment of the present invention, the device further comprises:
[0121] The information acquisition module is used to obtain the status and parameters set by the user for various functional modules, wherein the functional modules include a gesture control functional module, an element replacement functional module, and a device control functional module.
[0122] Furthermore, in one embodiment of the present invention, when it is obtained that the user has set the gesture control function module to be enabled, the image generation module includes:
[0123] A first gesture recognition unit is used to recognize whether there is preset gesture information in the collected image information;
[0124] a first image generating unit configured to load a preset number of function module setting buttons into the image information and perform image processing to generate a binocular image when the first gesture recognition unit recognizes that the collected image information contains preset gesture information;
[0125] The second image generating unit is configured to directly perform image processing based on the image information to generate a binocular image when the gesture recognition unit recognizes that the preset gesture information does not exist in the collected image information.
[0126] Furthermore, in one embodiment of the present invention, when it is obtained that the user has set the element replacement function module to be enabled, the image generation module includes:
[0127] A feature information recognition unit, used to recognize feature information of each person contained in the collected image information;
[0128] The third image generation unit is used to add or synthesize the human features into virtual elements according to the parameters set by the user for the element replacement function module, and perform image processing to generate a binocular image.
[0129] Furthermore, in one embodiment of the present invention, the image generation module further includes:
[0130] a second gesture recognition unit, configured to recognize whether there is switching gesture information in the collected image information;
[0131] The element switching unit is configured to switch the virtual element accordingly according to the switching gesture when the second gesture recognition unit recognizes that the collected image information contains switching gesture information.
[0132] Furthermore, in one embodiment of the present invention, when it is obtained that the user has set the device control function module to be enabled, the apparatus further includes:
[0133] a third gesture recognition unit, configured to recognize whether control gesture information exists in the collected image information;
[0134] The device control unit is configured to control the operation of the connected device according to the parameters set by the user for the device control function module when the third gesture recognition unit recognizes that the collected image information contains control gesture information.
[0135] Furthermore, in one embodiment of the present invention, the image generation module further includes:
[0136] An information loading unit, configured to load the state and parameters corresponding to the clicked function module upon recognizing the user's click gesture information for a setting button of any function module;
[0137] The information modification unit is used to modify the functional module according to the status and parameters set by the user.
[0138] Furthermore, in one embodiment of the present invention, the image generation module includes:
[0139] The fourth image generation unit is used to upload the currently collected image information to the cloud server in real time and synchronously, and receive the binocular image returned by the cloud server after performing real-time image processing based on the image information;
[0140] The fifth image generating unit is configured to perform local image calculation processing on the currently acquired image information to generate a binocular image.
[0141] Furthermore, in one embodiment of the present invention, the image generation module includes:
[0142] a sixth image generating unit, configured to perform image processing on two sets of images captured by the binocular camera at different viewing angles to generate a binocular image;
[0143] The seventh image generation unit is used to process the single set of images collected by the monocular camera to generate a virtual right image, and to compare the single set of images with the virtual right image. Figure 1 Perform image processing to generate binocular images;
[0144] The eighth image generating unit is used to perform image processing on a single set of images captured by the monocular camera to generate a binocular image.
[0145] The image processing device provided in the embodiment of the present invention has the same implementation principle and technical effects as those of the aforementioned method embodiment. For the sake of brief description, for matters not mentioned in the device embodiment, reference can be made to the corresponding content in the aforementioned method embodiment.
[0146] Example 4
[0147] Another aspect of the present invention also provides an intelligent terminal, see Figure 3 , shown is a smart terminal in the fourth embodiment of the present invention, including a memory 20, a processor 10, and a program 30 stored in the memory and executable on the processor. When the processor 10 executes the program 30, the image processing method described above is implemented.
[0148] In some embodiments, the processor 10 can be a central processing unit (CPU), a controller, a microcontroller, a microprocessor or other data processing chip, used to run the program code stored in the memory 20 or process data, such as executing access restriction programs.
[0149] The memory 20 includes at least one type of readable storage medium, including a flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 20 can be an internal storage unit of the smart terminal, such as the hard disk of the smart terminal. In other embodiments, the memory 20 can also be an external storage device of the smart terminal, such as a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card, etc. equipped on the smart terminal. Furthermore, the memory 20 can also include both an internal storage unit of the smart terminal and an external storage device. The memory 20 can be used not only to store application software and various types of data installed in the smart terminal, but also to temporarily store data that has been output or is to be output.
[0150] It should be pointed out that Figure 3 The structure shown does not constitute a limitation on the smart terminal. In other embodiments, the smart terminal may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0151] In summary, the smart terminal in the above-mentioned embodiment of the present invention collects image information of the current environment in real time, and processes the currently collected image information to generate a binocular image. The currently generated binocular image is then loaded and displayed, that is, the binocular image is displayed on the screen of the smart terminal. After the smart terminal is loaded into the VR box, when the user uses the VR box to watch the binocular view, a three-dimensional stereo image will be synthesized in the human brain, thereby achieving a three-dimensional stereo effect, thereby achieving a real-time interactive experience for the user, avoiding the problem that the simple VR box composed of the existing simple smart terminal combined with the VR box requires a dedicated film source for the interactive experience. At the same time, the picture collected in real time by the smart terminal can be processed to produce a three-dimensional stereo effect, making the interaction easier. At the same time, since the user's own smart terminal is used as a carrier of the interactive experience, the user's usage cost can be effectively saved, solving the problem of high cost of existing MR equipment and VR equipment.
[0152] An embodiment of the present invention further provides a readable storage medium having a program stored thereon, which implements the above-mentioned image processing method when executed by a processor.
[0153] Those skilled in the art will appreciate that the logic and / or steps represented in the flowcharts or otherwise described herein, for example, may be considered as a sequenced list of executable instructions for implementing the logical functions, and may be embodied in any readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "readable storage medium" may be any device that can contain, store, communicate, propagate, or transmit a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.
[0154] More specific examples (a non-exhaustive list) of readable storage media include the following: an electrical connection with one or more wires (electronic device), a portable computer disk cartridge (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the readable storage medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering, or processing in another suitable manner as necessary, and then stored in a memory.
[0155] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement the hardware: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0156] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "examples," "specific examples," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, schematic representations of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.
[0157] The above-described embodiments merely represent several implementation methods of the present invention. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art would be able to make various modifications and improvements without departing from the spirit of the present invention, all of which fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention shall be determined by the appended claims.
Claims
1. An image processing method, applied to a smart terminal, characterized in that: The method comprises: Obtaining the status and parameters set by the user for various functional modules, including the gesture control functional module, the element replacement functional module, and the device control functional module; Real-time collection of image information in the current environment; Performing image processing on the currently acquired image information to generate a binocular image, wherein the binocular image is a binocular view image in which the left view and the right view are located on the left and right sides respectively; Load and display the currently generated binocular image; When it is obtained that the user sets the gesture control function module to an enabled state, the step of performing image processing based on the currently collected image information to generate a binocular image includes: Identify whether there is preset gesture information in the collected image information; If yes, a preset number of function module setting buttons are loaded into the image information, and image processing is performed to generate a binocular image; If not, directly perform image processing based on the image information to generate a binocular image; After the step of loading a preset number of function module setting buttons into the image information, the method further includes: When a user clicks a button on any function module, the state and parameters of the clicked function module are loaded; Modify the functional modules accordingly according to the status and parameters set by the user.
2. The image processing method according to claim 1, wherein: When it is obtained that the user sets the element replacement function module to be enabled, the step of performing image processing based on the currently collected image information to generate a binocular image includes: Identify the characteristic information of each person contained in the collected image information; According to the parameters set by the user for the element replacement function module, the human features are added or synthesized and replaced with virtual elements, and image processing is performed to generate a binocular image.
3. The image processing method according to claim 2, wherein: After the step of adding or synthesizing the character features into virtual elements according to the parameters set by the user for the element replacement function module, the method further includes: Identifying whether there is switching gesture information in the collected image information; If so, the virtual element is switched accordingly according to the switching gesture.
4. The image processing method according to claim 1, wherein: When the user obtains the device control function module When set to the enabled state, the method further includes: Identifying whether there is control gesture information in the collected image information; If so, the operation of the connected device is controlled accordingly according to the parameters set by the user for the device control function module.
5. The image processing method according to claim 1, wherein: The step of performing image processing to generate a binocular image according to the currently collected image information includes: Upload the currently collected image information to the cloud server in real time and receive the binocular image returned by the cloud server after real-time image processing based on the image information; or The currently collected image information is processed locally to generate a binocular image.
6. The image processing method according to claim 1, wherein: The step of performing image processing to generate a binocular image according to the currently collected image information includes: Performing image processing on two sets of images captured by a binocular camera at different viewing angles to generate a binocular image; or Processing a single set of images captured by a monocular camera to generate a virtual right image, and processing the single set of images together with the virtual right image to generate a binocular image; or The single set of images captured by the monocular camera is processed to generate a binocular image.
7. An intelligent terminal, characterized in that: The image processing method comprises a memory, a processor, and a program stored in the memory and executable on the processor, wherein the processor implements the image processing method according to any one of claims 1 to 6 when executing the program.
Citation Information
Patent Citations
Method for displaying real object in head-mounted display, and head-mounted display for displaying real object
CN106484085A
AR display method and system based on binocular camera and VR equipment
CN112114667A