Image processing method, device, electronic equipment and program product

By directly identifying target objects in real-world images and displaying associated display objects, the high computational resource consumption problem in existing augmented reality technologies is solved, achieving more efficient display and a better user experience.

CN114863305BActive Publication Date: 2026-05-15FACE CUTE CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FACE CUTE CO LTD
Filing Date
2021-02-03
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing augmented reality technology requires a large amount of computing resources to display virtual objects, resulting in low display efficiency and poor universality.

Method used

By directly utilizing image recognition technology at the image level, the target object image is determined from the real-world captured image, and the display object associated with the target object is obtained. The display position is directly determined and displayed in the real-world captured image, avoiding the need to rebuild the virtual model.

Benefits of technology

It saves computing resources, improves display efficiency, and enhances the user's interactive and visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114863305B_ABST
    Figure CN114863305B_ABST
Patent Text Reader

Abstract

The image processing method, device, electronic device and program product provided by the embodiments of the present disclosure obtain a real scene shooting image; determine a target object image in the real scene shooting image, the target object image being an image including a target object; acquire a display object associated with the target object; determine a target display position of the display object in the real scene shooting image according to the target object image; and display the display object at the target display position in the real scene shooting image; the technical solution provided by the embodiments makes it possible to directly display a display object associated with a target object by using an augmented reality display technology, and since it is not necessary to reconstruct a virtual model, the calculation resources are saved, the display efficiency is improved, and the user can obtain better interactive experience and visual experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computers, and more particularly to an image processing method, apparatus, electronic device, and program product. Background Technology

[0002] Augmented Reality (AR) is a technology that cleverly integrates virtual information with the real world.

[0003] Augmented reality (AR) has emerged as a potential method for information presentation. In existing technologies, displaying virtual objects associated with a real-world object using AR is a common practice. However, displaying these virtual objects often requires their reconstruction.

[0004] Obviously, this method requires a lot of computing resources to display virtual objects, resulting in low display efficiency and poor universality. Summary of the Invention

[0005] To address the aforementioned issues, this disclosure provides an image processing method, apparatus, electronic device, and program product.

[0006] In a first aspect, embodiments of this disclosure provide an image processing method, including:

[0007] Obtain real-scene images;

[0008] The target object image is determined from the real-scene captured image, wherein the target object image is an image that includes the target object;

[0009] Obtain the display object associated with the target object;

[0010] Based on the image of the target object, determine the target display position of the display object in the real-scene captured image; and

[0011] The display object is displayed at the target display position in the real-scene image.

[0012] In a second aspect, embodiments of this disclosure provide an image processing apparatus, comprising:

[0013] The image capture and display module is used to obtain real-scene images.

[0014] The recognition processing module is used to determine a target object image in the real-scene captured image, the target object image including the target object; it is also used to obtain a display object associated with the target object; and to determine the target display position of the display object in the real-scene captured image based on the target object image.

[0015] The shooting display module is also used to display the display object at the target display position in the real-scene shooting image.

[0016] Thirdly, embodiments of this disclosure provide an electronic device, including: at least one processor and a memory;

[0017] The memory stores computer-executed instructions;

[0018] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the first aspect above and various possible methods relating to the image processing described in the first aspect.

[0019] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the image processing method described in the first aspect and various possible designs of the first aspect.

[0020] Fifthly, embodiments of this disclosure provide a computer program product including computer instructions that, when executed by a processor, provide the image processing method as described in the first aspect and various possible designs of the first aspect.

[0021] The image processing method, apparatus, electronic device, and program product provided in this disclosure obtain a real-scene captured image; determine a target object image in the real-scene captured image, wherein the target object image is an image including the target object; acquire a display object associated with the target object; determine the target display position of the display object in the real-scene captured image based on the target object image; and display the display object at the target display position in the real-scene captured image. The technical solution provided in this embodiment enables the direct display of a display object associated with a target object using augmented reality display technology. Since it is not necessary to reconstruct a virtual model, it saves computing resources, improves display efficiency, and allows users to obtain a better interactive and visual experience. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram of a network architecture on which this disclosure is based;

[0024] Figure 2 A schematic flowchart of an image processing method provided in an embodiment of this disclosure;

[0025] Figure 3 A schematic diagram of a first interface of an image processing method provided in a disclosed embodiment;

[0026] Figure 4 A schematic diagram of a second interface of an image processing method provided in a disclosed embodiment;

[0027] Figure 5 A schematic diagram of a third interface of an image processing method provided in a disclosed embodiment;

[0028] Figure 6 A schematic diagram of a fourth interface of an image processing method provided in a disclosed embodiment;

[0029] Figure 7 A schematic diagram of the fifth interface of an image processing method provided in a disclosed embodiment;

[0030] Figure 8 A schematic diagram of the sixth interface of an image processing method provided in a disclosed embodiment;

[0031] Figure 9 This is a structural block diagram of an image processing apparatus provided in an embodiment of the present disclosure;

[0032] Figure 10 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0033] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0034] Augmented Reality (AR) is a technology that cleverly integrates virtual information with the real world.

[0035] Augmented reality (AR) has emerged as a potential method for information presentation. In existing technologies, when displaying AR, the terminal first captures a real-world scene to obtain an image. Then, using AR technology, virtual objects are overlaid onto the captured image, and the resulting image is presented to the user.

[0036] When overlaying these virtual objects, it is first necessary to construct related virtual objects based on the objects in the real scene, and then display the objects on the real scene image based on the constructed virtual objects.

[0037] In existing technologies, the construction of virtual objects is generally achieved through Simultaneous Localization and Mapping (SLAM) technology. Specifically, the terminal first scans the environment in real time to construct environment-based virtual objects (such as 3D virtual models), then loads the display objects to be shown into the virtual objects, and finally overlays the virtual objects loaded with display objects onto the real scene to complete the display.

[0038] However, augmented reality display technology based on SLAM needs to regenerate and reconstruct virtual objects, including 3D virtual models, every time it is displayed. The generation and construction processes require a lot of computing resources and time costs, which makes the display time of SLAM-based augmented reality display technology longer and the display efficiency lower.

[0039] To address this problem, the inventors, through research, creatively discovered that a low-cost method can be used when presenting display objects associated with objects. According to the method of this disclosure, image recognition technology at the image layer level can be directly used to display display objects in real-world images.

[0040] Specifically, the process involves: obtaining a real-scene captured image; determining a target object image within the real-scene captured image, wherein the target object image is an image including the target object; acquiring a display object associated with the target object; determining the target display position of the display object in the real-scene captured image based on the target object image; and displaying the display object at the target display position in the real-scene captured image.

[0041] The technical solution provided in this embodiment enables the direct display of display objects associated with the target object using augmented reality display technology. Since it is not necessary to rebuild the virtual model, it saves computing resources, improves display efficiency, and allows users to have a better interactive and visual experience.

[0042] refer to Figure 1 , Figure 1 This is a schematic diagram of a network architecture upon which this disclosure is based. Figure 1 The network architecture shown may specifically include terminal 1 and server 2.

[0043] Specifically, terminal 1 can be a user's mobile phone, smart home device, tablet computer, wearable device, or other hardware device that can capture and display real-world scenes. Terminal 1 can integrate or install an image processing device, which is hardware or software used to execute the image processing method disclosed herein. The image processing device can provide terminal 1 with an augmented reality display page, and terminal 1 uses its screen or display components to display the augmented reality display page provided by the image processing device to the user.

[0044] Server 2 may specifically be a server or server cluster located in the cloud, and the server or server cluster may store image data and display object data related to the image processing method provided in this disclosure.

[0045] Specifically, when performing the image processing method provided in this disclosure, the image processing device can also use the network component of terminal 1 to interact with server 2, obtain image data and display object data stored in server 2, and perform corresponding processing and display.

[0046] Figure 1 The architecture shown is applicable to the field of information presentation; in other words, it can be used to present display objects associated with objects in real-world scenarios across various environments.

[0047] For example, the image processing method provided in this disclosure can be applied to scenarios based on augmented reality displays. For instance, in scenarios where landmark buildings are combined with augmented reality display technology to display "New Year (Spring Festival) greetings" or "supporting a certain place," the "landmark buildings" in the real-world scene can be identified first. Then, display objects associated with the "landmark buildings," including "New Year (Spring Festival) greetings" or "supporting a certain place," can be displayed in the real-world scene including the "landmark buildings." In this scenario, the display of display objects associated with the "landmark buildings," including "New Year (Spring Festival) greetings" or "supporting a certain place," can be achieved using the image processing method provided in this disclosure.

[0048] For example, in some "treasure hunt" games based on augmented reality display technology, the image processing method provided in this disclosure can be used to display the "treasure chest" as a display object associated with clue items in the real scene in the "treasure hunt" game; or, in some "farmland simulation" games based on augmented reality display technology, the image processing method provided in this disclosure can be used to display the "farmland" as a display object associated with farmland markers in the real scene in the "farmland simulation" game.

[0049] For example, the image processing method provided in this disclosure can also be applied to advertising scenarios based on augmented reality displays. For instance, for some goods or products, the image processing method provided in this disclosure can be used to present related product reviews or product introductions, thereby providing users with more information about the goods and improving the user experience.

[0050] In addition, in some everyday scenarios where the camera function is used, the image processing method provided in this disclosure can also be used to display objects, thereby presenting users with more information about the scene and increasing the user's interactive experience.

[0051] To give a further example, in scenarios that require turning on the terminal camera to take real-world photos, such as "scanning a QR code to make a payment" or "taking a picture," the image processing method provided in this disclosure can be executed simultaneously to present information in various daily life scenarios.

[0052] The image processing method provided in this disclosure will be further explained below:

[0053] Firstly, Figure 2 This is a schematic flowchart illustrating an image processing method provided in an embodiment of this disclosure. (Reference) Figure 2 The image processing method provided in this disclosure includes:

[0054] Step 101: Obtain real-scene images;

[0055] Step 102: Determine the target object image from the real-scene captured image, wherein the target object image is an image that includes the target object;

[0056] Step 103: Obtain the display object associated with the target object;

[0057] Step 104: Determine the target display position of the display object in the real-scene captured image based on the target object image;

[0058] Step 105: Display the object at the target display position in the real-scene captured image.

[0059] It should be noted that the execution subject of the processing method provided in this embodiment is the aforementioned image processing device. In some embodiments of this disclosure, it specifically refers to a client or display terminal that can be installed or integrated on a terminal. Users can operate the image processing device through the terminal, so that the image processing device can respond to the operations triggered by the user.

[0060] First, the terminal will obtain a real-scene image. This real-scene image can be an image obtained by the terminal calling its own shooting component to capture the current environment, or it can be a real-time image of the real scene obtained by the image processing device through other means.

[0061] Subsequently, the terminal will identify the target object in the real-scene captured image to determine whether there is a target object in the real-scene captured image that can be used to perform the image processing method provided in this disclosure, that is, to determine whether there is an image containing a target object in the real-scene captured image.

[0062] Figure 3 This is a schematic diagram of the first interface of an image processing method provided in a disclosed embodiment. In a real-world scenario, the target object mentioned in this disclosure can be any object appearing in the real-world scene, including but not limited to "landmark buildings", "goods" and "activity QR codes" in the aforementioned scenario.

[0063] like Figure 3 As shown, the target object here can be a "beverage bottle". By performing recognition processing on the real-scene captured image 300, the target object image 301 containing the target object "beverage bottle" in the real-scene captured image 300 is determined.

[0064] In an optional implementation, when determining a target object, the terminal can perform feature matching by comparing the features of the currently captured real-scene image with the features of a reference image of the target object.

[0065] It should be noted that the reference image is generally an image containing the target object that is pre-set in the image database of the server. The reference image generally includes a standard image of the target object, such as a front view of a landmark building, or a front view of a product, or an image including all QR code information, etc.

[0066] Figure 4 This is a schematic diagram of a second interface of an image processing method provided in a disclosed embodiment, with reference to... Figure 4 To determine the target object, it is first necessary to acquire an image containing the target object, namely reference image 302. Then, the terminal can determine the target object image 301 in the real-scene captured image 300 based on the reference image 302.

[0067] The image database of a server often stores a large number of reference images of objects, all of which are pre-stored in the image database.

[0068] In the embodiments provided in this disclosure, in order to identify the target object in the real-scene image, it is first necessary to find the image corresponding to the current real-scene image from a large number of reference images stored in the image database as a reference image.

[0069] Optionally, this process can be implemented based on feature matching: First, extract the global features of the real-scene image; then, match the global features of the real-scene image with the global features of at least one reference image stored in the image database; finally, determine the reference image whose features match those of the real-scene image as the reference image. This embodiment of the present disclosure, by extracting the global features of the real-scene image and using these global features for matching to determine the reference image, helps to improve the matching degree with the real-scene image and enhances the user experience.

[0070] Specifically, the extraction of global features can be achieved using existing machine learning models, such as convolutional neural networks. To identify, as far as possible, whether there are target objects in the current real-world image that can be used to execute the scheme disclosed herein, global features representing global image information or global characteristics can be extracted from the real-world image during identification for subsequent matching.

[0071] Subsequently, in the process of matching the extracted global features of the real-scene image with the global features of at least one reference image stored in the image database, to improve matching efficiency, the global features of each reference image pre-extracted and stored in the image database can be pre-extracted and stored. After obtaining the global features of the real-scene image, a one-to-one matching process can be performed between these global features and the pre-stored global features of the reference images. This matching process includes, but is not limited to, feature comparison processing and image similarity calculation based on the distance between features.

[0072] This matching process allows us to find a reference image from the image database that has the most similar global features to the currently captured real-world image, which will then be used as the reference image for this process.

[0073] In other alternative implementations, the above-mentioned acquisition of reference images can also be achieved through other more efficient methods, such as filtering each reference image in the server's image database based on the terminal's current location information to obtain the corresponding reference image.

[0074] Specifically, in some scenarios, since the target object is strongly associated with geographical location information, the terminal's current location information can also be used as a filtering condition for acquiring each reference image. That is, the terminal can upload its current location information to the server and receive the reference images associated with the current location information returned by the server.

[0075] Figure 5 A third interface diagram of an image processing method provided in the disclosed embodiments shows that the terminal can upload the current location information "a street garden in Beijing" to the server. (Refer to...) Figure 5 The server can determine the reference image 402 corresponding to the "farmland QR code" at the current location based on the current location information provided by the terminal. Then, the server sends the reference image 402 to the terminal for processing to determine the target object image 401 of the "farmland QR code" and display the "farmland" display object 403 corresponding to the "farmland QR code" on the terminal screen.

[0076] The aforementioned current location information may include geographic location information, such as geographic coordinates and latitude and longitude; it may also include directional location information, such as facing northwest at 45 degrees. During the process of the server determining the corresponding reference image from a pre-stored set of reference images based on the terminal's current location information, the server can push target objects that may exist in the currently captured real-world image by combining geographic location information and directional location information, and then send the corresponding reference image to the terminal.

[0077] In other alternative implementations or application scenarios, the two methods described above (feature-based matching and location-based matching) can be used simultaneously in the process of acquiring the reference image. For example, at least one reference image associated with the current location information returned by the server can be received first based on the terminal's current location information; then, based on the global features extracted from the real-scene captured image, a matching process is performed with the global features of the at least one reference image; finally, the reference image that matches the features of the real-scene captured image is determined as the reference image. The specific implementation method is similar to the principle in the aforementioned process and will not be elaborated here.

[0078] Of course, in other alternative embodiments, the target object identification process involved in this disclosure can be implemented based on existing recognition models, recognition components, or recognition servers that can be used to identify whether an image contains a target object. Specifically, the terminal can send the captured real-scene image to the recognition module, recognition component, or recognition server, which will then perform the recognition processing to determine whether the real-scene image contains a target object and return the processing result to the terminal.

[0079] Using the above method, images containing target objects, i.e., target object images, can be identified from real-scene captured images. After determining the target object image, the step of obtaining the display object associated with the target object will also be performed.

[0080] Among them, the display object refers to static or dynamic objects such as text information, image information, and virtual effects associated with the target object that can be presented on the terminal screen.

[0081] For example, in the aforementioned scenario of combining landmark buildings with augmented reality display technology to send "New Year (Spring Festival) blessings" or "support a certain place," the text information of the blessing or the image information of the support can serve as the display object associated with the landmark building; while in the aforementioned "treasure hunt" game, the virtual effects of the treasure chest can serve as the display object of a certain clue item in the real scene.

[0082] In an optional implementation, the display object can be obtained through interaction with the server. That is, after confirming the target object image, the terminal uploads the target object image to the server and receives the display object returned by the server to perform the display of the display object.

[0083] Specifically, for display objects associated with the target object, they are generally pre-set, that is, by establishing a mapping relationship between the display object and the target object in the server, so that when the server receives an image of the target object containing the target object, it can send the display object corresponding to the target object to the terminal based on the mapping relationship, so that the terminal can display the display object.

[0084] Furthermore, synchronously or asynchronously with the step of acquiring the display object, the terminal will also determine the target display position of the display object in the real-scene captured image based on the target object image, which can be specifically achieved through steps 1041 and 1042:

[0085] Step 1041: Perform model adaptation between the target object image and the reference image;

[0086] Step 1042: Based on the preset display position in the reference image, determine the target display position of the display object in the real-scene captured image.

[0087] The reference image includes a preset display position. Model matching is used to determine the target display position of the display object in the real-scene image, avoiding an abrupt appearance of the display object and improving its display effect.

[0088] When displaying an object, it is necessary not only to obtain the object's content, but also to determine the object's location within the captured image.

[0089] Based on this, when configuring associated display objects for a target object, a corresponding preset display position can also be configured for the display object; and in order to better perform image positioning, the preset display position can be carried in the reference image and acquired and processed by the terminal.

[0090] Specifically, Figure 6This is a schematic diagram of a fourth interface of an image processing method provided in a disclosed embodiment. As mentioned above, the reference image includes an image containing all information about the target object, wherein the combination... Figure 6 As shown, the reference image 602 will also include a preset display position 604, which will be used to represent the display position of the display object associated with the target object in the reference image 602. By performing model adaptation on the target object image 601 and the reference image 602, the target display position 605 of the display object in the real-scene captured image is determined based on the preset display position 604 in the reference image 602.

[0091] In particular, model adaptation for the image in step 1041 can be achieved through image feature matching:

[0092] First, the terminal can determine the affine transformation matrix of the image based on the reference image and the target object image.

[0093] There are various ways to obtain the affine transformation matrix. In one optional implementation, the affine transformation matrix can be obtained through image feature matching technology: First, extract the local features of the target object image; then, match the local features of the target object image with the local features of the reference image to obtain the mapping relationship between the local features; finally, determine the image affine transformation matrix based on the mapping relationship between the local features.

[0094] Similar to the aforementioned methods, the extraction of local features can be achieved using existing machine learning models, such as convolutional neural networks. However, unlike the previous methods, in determining the image radiometric transformation matrix, to identify the feature correspondence between the target object image and the reference image as accurately as possible, local features representing local information or characteristics of the image can be extracted from the target object image for subsequent matching.

[0095] In an optional implementation, the local features of both the target object image and the reference image can be robust SIFT (Distinctive Image Features from Scale-Invariant Keypoints) features, and the data dimension of the SIFT feature can be 128 dimensions.

[0096] In an optional implementation, the matching and mapping relationship between local features of the target object image and the reference image can be obtained by using the nearest neighbor search method based on brute-force / linear-scan. The distance between matching features can be selected based on the standard L2Eclidean Distance, which means that the distance of the current nearest matching feature point is smaller than the distance of the previous nearest matching feature point by a certain ratio.

[0097] In an optional implementation, after obtaining the above mapping relationship, an affine matrix between images is established based on this mapping relationship. That is, using the RanSac framework, the image affine transformation matrix is ​​obtained through random sampling and multiple iterations. It is known that the target object image can be obtained by performing a dot product operation between the reference image and this image affine transformation matrix.

[0098] Accordingly, in step 1042, the terminal performs an affine transformation on the preset display position in the reference image according to the image affine transformation matrix to obtain the target display position of the display object in the real-scene captured image.

[0099] To reduce computational complexity, the target display position of the object in a real-world photograph can be determined using reference coordinate points. Figure 7 A schematic diagram of the fifth interface of an image processing method provided in the disclosed embodiments, as shown below. Figure 7 As shown, the reference image may include a reference reference coordinate point 6041 with a preset display position.

[0100] First, an affine transformation can be performed on the reference reference coordinate point 6041 according to the image affine transformation matrix to obtain the transformed target reference coordinate point 6051. Then, based on the target reference coordinate point 6051, the target display position 605 of the display object 603 in the real-scene captured image can be determined. The reference reference coordinate point 6041, which includes the preset display position 604 in the reference image 602, is preset.

[0101] Finally, the terminal displays the object at the target display position in the real-scene captured image. The image processing device can use augmented reality display technology to overlay the object onto the target display position in the real-scene captured image. This facilitates seamless integration of the displayed object with the real scene, enhances the interactive experience, and increases the richness of the display.

[0102] In addition, in other alternative implementations, in order to provide users with a more interactive experience, the terminal can also respond to the user's trigger operation on the display object displayed in the real-scene captured image; and display the trigger result on the display object in the real-scene captured image.

[0103] In different scenarios, because the displayed objects are different, the triggering results will vary depending on the actual scenario:

[0104] For example, in a simulation game, after the "farmland" object is displayed in the real-world image, the user can trigger the "harvest agricultural products" action to display the "harvest completed" result on the screen; or the user can trigger the "visit friend's farmland" action to display the "friend's farmland" result on the screen.

[0105] The examples above are merely illustrative. Depending on the actual application scenarios and business needs, there are many more possible implementation methods based on this disclosure, and this disclosure does not impose any restrictions on them.

[0106] To further describe the image processing method provided in this disclosure, the following will take a scenario of combining landmark buildings with augmented reality display technology to send "New Year (Spring Festival) greetings" as an example to provide a detailed explanation of the solution provided in this disclosure. In this embodiment, the target object can be a landmark building.

[0107] Figure 8 A sixth interface diagram of an image processing method provided in a disclosed embodiment is shown below. Figure 8 As shown, the user can trigger the terminal's shooting function to enable the terminal to call the shooting component to obtain a real-scene shooting image 800. The terminal can determine the target object image 801 of the landmark building in the real-scene shooting image 800 based on the current location information.

[0108] Using the processing method provided by the above embodiments, the real-scene captured image 800 will display a display object 803 associated with the target object, namely "Click here to send New Year's greetings" in the figure.

[0109] Optional, see reference Figure 8 The real-scene image 800 can also display the operation result 806 of the trigger operation on the display object 803 in the real-scene image 800; optionally, the real-scene image 800 can also display the corresponding landmark name, such as Figure 8 Shanghai is shown.

[0110] The image processing method provided in this embodiment obtains a real-scene captured image; determines a target object image in the real-scene captured image, wherein the target object image is an image including the target object; acquires a display object associated with the target object; determines the target display position of the display object in the real-scene captured image based on the target object image; and displays the display object at the target display position in the real-scene captured image. The technical solution provided in this embodiment enables the direct display of the display object associated with the target object using augmented reality display technology. Since it is not necessary to reconstruct the virtual model, it saves computing resources, improves display efficiency, and allows users to obtain a better interactive and visual experience.

[0111] Corresponding to the image processing method in the above embodiments, Figure 9 This is a structural block diagram of an image processing apparatus provided according to embodiments of the present disclosure. For ease of explanation, only the parts relevant to embodiments of the present disclosure are shown. (Refer to...) Figure 9 The image processing device includes: a capture and display module 10 and a recognition and processing module 20;

[0112] The shooting display module 10 is used to obtain real-scene shooting images;

[0113] The recognition processing module 20 is configured to determine a target object image in the real-scene captured image, the target object image including the target object; and to acquire a display object associated with the target object; and to determine the target display position of the display object in the real-scene captured image based on the target object image.

[0114] The shooting display module 10 is also used to display the display object at the target display position in the real-scene shooting image.

[0115] In an optional embodiment, when the recognition processing module 20 determines the target object image in the real-scene captured image, it is specifically used to: acquire a reference image, wherein the reference image is an image including the target object; and determine the target object image in the real-scene captured image based on the reference image.

[0116] In an optional embodiment, when the recognition processing module 20 performs the acquisition of the reference image, it is specifically used to: extract the global features of the real-scene image; match the global features of the real-scene image with the global features of at least one reference image stored in the image database; and determine the reference image that matches the features of the real-scene image as the reference image.

[0117] In an optional embodiment, when the recognition processing module 20 performs the acquisition of the reference image, it is specifically used to: upload the current location information to the server; and receive the reference image associated with the current location information returned by the server.

[0118] In an optional embodiment, the reference image includes a preset display position;

[0119] When the recognition processing module 20 performs the step of determining the target display position of the display object in the real-scene captured image based on the target object image, it is specifically used to: perform model adaptation between the target object image and the reference image, so as to determine the target display position of the display object in the real-scene captured image based on the preset display position in the reference image.

[0120] In an optional embodiment, when the recognition processing module 20 performs model adaptation between the target object image and the reference image to determine the target display position of the display object in the real-scene captured image based on the preset display position in the reference image, it is specifically used to: determine the image affine transformation matrix according to the reference image and the target object image; and perform an affine transformation on the preset display position in the reference image according to the image affine transformation matrix to obtain the target display position of the display object in the real-scene captured image.

[0121] In an optional embodiment, when the recognition processing module 20 performs the step of determining the image affine transformation matrix based on the reference image and the target object image, it is specifically used for:

[0122] Local features of the target object image are extracted; the local features of the target object image are matched with the local features of the reference image to obtain the mapping relationship between the local features; the affine transformation matrix of the image is determined based on the mapping relationship between the local features.

[0123] In an optional embodiment, the reference image includes a reference coordinate point at a preset display position;

[0124] When the recognition processing module 20 performs the step of performing an affine transformation on a preset display position in the reference image based on the image affine transformation matrix to obtain the target display position of the display object in the real-scene captured image, it is specifically used for:

[0125] Based on the affine transformation matrix of the image, the reference reference coordinate point is subjected to an affine transformation to obtain the transformed target reference coordinate point; based on the target reference coordinate point, the target display position of the display object in the real-scene captured image is determined.

[0126] In an optional embodiment, the reference reference coordinate point in the reference image, which includes a preset display position, is pre-set.

[0127] In an optional embodiment, after the recognition processing module 20 determines the target object image in the real-scene captured image, it is further configured to upload the target object image to the server; and receive the display object returned by the server, so that the capture display module 10 can perform the display of the display object.

[0128] In an optional embodiment, when the shooting and display module 10 performs the action of displaying the display object at the target display position in the real-scene captured image, it is specifically used for:

[0129] Based on augmented reality display technology, the display object is superimposed on the target display position in the real-scene captured image.

[0130] In an optional embodiment, the shooting display module 10 is further configured to respond to a user's trigger operation on a display object displayed in the real-scene shooting image; and display the trigger result on the display object in the real-scene shooting image.

[0131] In an optional embodiment, the target object is a landmark building.

[0132] In an optional embodiment, the shooting display module 10 is further configured to display the landmark name of the landmark building in the real-scene shooting image.

[0133] The image processing apparatus provided in this embodiment obtains a real-scene captured image; determines a target object image in the real-scene captured image, the target object image being an image including the target object; acquires a display object associated with the target object; determines a target display position of the display object in the real-scene captured image based on the target object image; and displays the display object at the target display position in the real-scene captured image. The technical solution provided in this embodiment enables the direct display of a display object associated with a target object using augmented reality display technology. Since it is not necessary to reconstruct a virtual model, this saves computing resources, improves display efficiency, and provides users with a better interactive and visual experience.

[0134] The electronic device provided in this embodiment can be used to execute the technical solutions of the above method embodiments. Its implementation principle and technical effects are similar, and will not be described again here.

[0135] refer to Figure 10The diagram illustrates a structural schematic of an electronic device 900 suitable for implementing embodiments of the present disclosure. The electronic device 900 can be a terminal device or a media library. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. Figure 10 The electronic device shown is merely one embodiment and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0136] like Figure 10 As shown, the electronic device 900 may include a processor 901 for executing image processing methods (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 902 or a program loaded from storage device 908 into random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The image processing method 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0137] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 10 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0138] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts according to embodiments of this disclosure. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by the image processing method 901, it performs the functions defined in the methods of embodiments of this disclosure.

[0139] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0140] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0141] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0142] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or media library. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0143] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0144] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0145] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0146] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific embodiments of machine-readable storage media may include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0147] The following are some embodiments of this disclosure.

[0148] In a first aspect, according to one or more embodiments of the present disclosure, an image processing method includes:

[0149] Obtain real-scene images;

[0150] The target object image is determined from the real-scene captured image, and the target object image is an image that includes the target object;

[0151] Obtain the display object associated with the target object;

[0152] Based on the target object image, determine the target display position of the display object in the real-scene captured image; and

[0153] The display object is displayed at the target display position in the real-scene captured image.

[0154] In an optional embodiment, determining the target object image from the real-scene captured image includes:

[0155] Acquire a reference image, wherein the reference image is an image including the target object;

[0156] The target object image is determined from the real-scene captured image based on the reference image.

[0157] In an optional embodiment, acquiring the reference image includes:

[0158] Extract global features from the real-scene captured images;

[0159] The global features of the real-scene captured image are matched with the global features of at least one reference image stored in the image database.

[0160] The reference image is determined by matching the features of the real-scene captured image.

[0161] In an optional embodiment, acquiring the reference image includes:

[0162] Upload the current location information to the server;

[0163] Receive a reference image returned by the server that is associated with the current location information.

[0164] In an optional embodiment, the reference image includes a preset display position;

[0165] Determining the target display position of the display object in the real-scene captured image based on the target object image includes:

[0166] The target object image is model-fitted with the reference image to determine the target display position of the display object in the real-scene captured image based on the preset display position in the reference image.

[0167] In an optional embodiment, the step of performing model adaptation between the target object image and the reference image to determine the target display position of the display object in the real-scene captured image based on a preset display position in the reference image includes:

[0168] Based on the reference image and the target object image, determine the image affine transformation matrix;

[0169] Based on the affine transformation matrix of the image, an affine transformation is performed on the preset display position in the reference image to obtain the target display position of the display object in the real-scene captured image.

[0170] In an optional embodiment, determining the image affine transformation matrix based on the reference image and the target object image includes:

[0171] Extract local features from the image of the target object;

[0172] The local features of the target object image are matched with the local features of the reference image to obtain the mapping relationship between the local features;

[0173] The image affine transformation matrix is ​​determined based on the mapping relationship between the local features.

[0174] In an optional embodiment, the reference image includes a reference coordinate point at a preset display position;

[0175] The step of performing an affine transformation on a preset display position in the reference image based on the image affine transformation matrix to obtain the target display position of the display object in the real-scene captured image includes:

[0176] Based on the image affine transformation matrix, the reference reference coordinate point is subjected to an affine transformation to obtain the transformed target reference coordinate point;

[0177] Based on the target reference coordinates, determine the target display position of the display object in the real-scene captured image.

[0178] In an optional embodiment, the reference reference coordinate point in the reference image, which includes a preset display position, is pre-set.

[0179] In an optional embodiment, after determining the target object image from the real-scene captured image, the method further includes:

[0180] Upload the image of the target object to the server;

[0181] Receive the display object returned by the server to perform the display of the display object.

[0182] In an optional embodiment, displaying the display object at the target display position in the real-scene captured image includes:

[0183] Based on augmented reality display technology, the display object is superimposed on the target display position in the real-scene captured image.

[0184] In an optional embodiment, the method further includes:

[0185] Responding to user trigger operations on display objects shown in the real-scene captured image;

[0186] The triggering result of the display object will be displayed in the real-scene captured image.

[0187] In an optional embodiment, the target object is a landmark building.

[0188] In an optional embodiment, the method further includes:

[0189] The landmark name of the landmark building is displayed in the real-scene image.

[0190] Secondly, according to one or more embodiments of the present disclosure, an image processing apparatus includes:

[0191] The image capture and display module is used to obtain real-scene images.

[0192] The recognition processing module is used to determine a target object image in the real-scene captured image, the target object image including the target object; it is also used to obtain a display object associated with the target object; and to determine the target display position of the display object in the real-scene captured image based on the target object image.

[0193] The shooting display module is also used to display the display object at the target display position in the real-scene shooting image.

[0194] In an optional embodiment, when the recognition processing module determines the target object image in the real-scene captured image, it is specifically used to: acquire a reference image, wherein the reference image is an image including the target object; and determine the target object image in the real-scene captured image based on the reference image.

[0195] In an optional embodiment, when the recognition processing module performs the acquisition of the reference image, it is specifically used to: extract the global features of the real-scene image; match the global features of the real-scene image with the global features of at least one reference image stored in the image database; and determine the reference image that matches the features of the real-scene image as the reference image.

[0196] In an optional embodiment, when the recognition processing module performs the acquisition of the reference image, it is specifically used to: upload the current location information to the server; and receive the reference image associated with the current location information returned by the server.

[0197] In an optional embodiment, the reference image includes a preset display position;

[0198] When the recognition processing module performs the step of determining the target display position of the display object in the real-scene captured image based on the target object image, it is specifically used to: perform model adaptation between the target object image and the reference image, so as to determine the target display position of the display object in the real-scene captured image based on the preset display position in the reference image.

[0199] In an optional embodiment, when the recognition processing module performs model adaptation between the target object image and the reference image to determine the target display position of the display object in the real-scene captured image based on the preset display position in the reference image, it is specifically used to: determine the image affine transformation matrix according to the reference image and the target object image; and perform an affine transformation on the preset display position in the reference image according to the image affine transformation matrix to obtain the target display position of the display object in the real-scene captured image.

[0200] In an optional embodiment, when the recognition processing module performs the step of determining the image affine transformation matrix based on the reference image and the target object image, it is specifically used for:

[0201] Local features of the target object image are extracted; the local features of the target object image are matched with the local features of the reference image to obtain the mapping relationship between the local features; the affine transformation matrix of the image is determined based on the mapping relationship between the local features.

[0202] In an optional embodiment, the reference image includes a reference coordinate point at a preset display position;

[0203] When the recognition processing module performs the affine transformation on the preset display position in the reference image based on the image affine transformation matrix to obtain the target display position of the display object in the real-scene captured image, it is specifically used for:

[0204] Based on the affine transformation matrix of the image, the reference reference coordinate point is subjected to an affine transformation to obtain the transformed target reference coordinate point; based on the target reference coordinate point, the target display position of the display object in the real-scene captured image is determined.

[0205] In an optional embodiment, the reference reference coordinate point in the reference image, which includes a preset display position, is pre-set.

[0206] In an optional embodiment, after the recognition processing module determines the target object image in the real-scene captured image, it is further configured to upload the target object image to the server; and receive the display object returned by the server, so that the capture display module can perform the display of the display object.

[0207] In an optional embodiment, when the shooting and display module performs the action of displaying the display object at the target display position in the real-scene captured image, it is specifically used for:

[0208] Based on augmented reality display technology, the display object is superimposed on the target display position in the real-scene captured image.

[0209] In an optional embodiment, the shooting display module is further configured to respond to a user's trigger operation on a display object displayed in the real-scene shooting image; and display the trigger result on the display object in the real-scene shooting image.

[0210] In an optional embodiment, the target object is a landmark building.

[0211] In an optional embodiment, the shooting display module is further configured to display the landmark name of the landmark building in the real-scene shooting image.

[0212] Thirdly, according to one or more embodiments of this disclosure, an electronic device includes:

[0213] At least one processor; and

[0214] Memory;

[0215] The memory stores computer-executed instructions;

[0216] The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in the preceding one.

[0217] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium stores computer-executable instructions that, when executed by a processor, implement the method as described in any of the preceding claims.

[0218] Fifthly, according to one or more embodiments of the present disclosure, a computer program product includes computer instructions that, when executed by a processor, implement the method as described in any of the preceding claims.

[0219] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0220] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0221] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely embodiments of the implementation of the claims.

Claims

1. An image processing method, characterized in that, include: Obtain real-scene images; The target object image is determined from the real-scene captured image, and the target object image is an image that includes the target object; Obtain the display object associated with the target object; The affine transformation matrix of the image is determined based on the mapping relationship obtained by matching the local features of the target object image with the local features of the reference image; the reference image includes a reference coordinate point at a preset display position. The preset display position is used to indicate the display position of the display object associated with the target object in the reference image; Based on the image affine transformation matrix, the reference reference coordinate point is subjected to an affine transformation to obtain the transformed target reference coordinate point; Based on the target reference coordinates, determine the target display position of the display object in the real-scene captured image; The display object is displayed at the target display position in the real-scene captured image.

2. The image processing method according to claim 1, characterized in that, Determining the target object image from the real-scene captured image includes: Acquire a reference image, wherein the reference image is an image including the target object; The target object image is determined from the real-scene captured image based on the reference image.

3. The image processing method according to claim 2, characterized in that, The acquisition of the reference image includes: Extract global features from the real-scene captured images; The global features of the real-scene captured image are matched with the global features of at least one reference image stored in the image database. The reference image is determined by matching the features of the real-scene captured image.

4. The image processing method according to claim 2, characterized in that, The acquisition of the reference image includes: Upload the current location information to the server; Receive a reference image returned by the server that is associated with the current location information.

5. The image processing method according to claim 1, characterized in that, The reference image includes a preset reference coordinate point at a pre-defined display position.

6. The image processing method according to claim 1, characterized in that, After determining the target object image from the real-scene captured image, the process further includes: Upload the image of the target object to the server; Receive the display object returned by the server to perform the display of the display object.

7. The image processing method according to claim 1, characterized in that, The step of displaying the display object at the target display position in the real-scene captured image includes: Based on augmented reality display technology, the display object is superimposed on the target display position in the real-scene captured image.

8. The image processing method according to claim 1, characterized in that, Also includes: Responding to user trigger operations on display objects shown in the real-scene captured image; The triggering result of the display object will be displayed in the real-scene captured image.

9. The image processing method according to claim 1, characterized in that, The target object is a landmark building.

10. The image processing method according to claim 9, characterized in that, Also includes: The landmark name of the landmark building is displayed in the real-scene image.

11. An image processing apparatus, characterized in that, include: The image capture and display module is used to obtain real-scene images. The recognition processing module is used to determine a target object image in the real-scene captured image, the target object image including the target object; it is also used to obtain a display object associated with the target object; and to determine an image affine transformation matrix based on a mapping relationship obtained by matching local features of the target object image with local features of a reference image; the reference image includes a reference coordinate point with a preset display position. The preset display position is used to indicate the display position of the display object associated with the target object in the reference image; Based on the image affine transformation matrix, the reference reference coordinate point is subjected to an affine transformation to obtain the transformed target reference coordinate point; Based on the target reference coordinate point, the target display position of the display object in the real-scene captured image is determined; the capture display module is also used to display the display object at the target display position in the real-scene captured image.

12. An electronic device, wherein, include: At least one processor; as well as Memory; The memory stores computer-executed instructions; The at least one processor executes computer execution instructions stored in the memory, causing the at least one processor to perform the method as described in any one of claims 1-10.

13. A computer-readable storage medium, wherein, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method as described in any one of claims 1-10.

14. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the method described in any one of claims 1-10.