A video implementation method, device, electronic device, and readable storage medium
By generating and adjusting the global image in AR remote guidance, the problem of unavailability of video calls in an unstable environment is solved, and the stability of video calls under unstable network conditions is achieved.
Patent Information
- Application Number
- CN202210494123.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-05
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2042-05-05
AI Technical Summary
In the AR remote guidance scenario, the network is unstable in the environment where the user is located, resulting in unavailable video calls.
By acquiring multiple images of different shooting angles sent by the first electronic device, a global image of the target object is generated, and the size and display area of the global image are adjusted according to the first posture information, so as to achieve stability of video calls.
Even if the network is unstable in the environment in which the first electronic device is located, the normal effect of video calls can be achieved, and the user can know the location of the target object that the other party is watching in real time.
Smart Images

Figure CN115103148B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to a video implementation method, apparatus, electronic device, and readable storage medium. Background Art
[0002] In the scenario of AR remote guidance, usually the user side conducts a video with the expert side through AR glasses, so as to achieve remote guidance of the user by the expert. However, if the network in the environment where the user side is located is unstable, such as weak network, unstable bandwidth, etc., it will directly cause the video call to be unavailable. Summary of the Invention
[0003] In view of this, embodiments of this application provide a video implementation method, apparatus, electronic device, and readable storage medium to at least solve the above technical problems existing in the prior art.
[0004] According to a first aspect of this application, embodiments of this application provide a video implementation method, including: obtaining multiple first images sent by a first electronic device, each first image being an image of a target object collected by the first electronic device, and the shooting angles of the target object in the multiple first images being different; generating a global image of the target object according to the multiple first images; obtaining first pose information sent by the first electronic device; and adjusting the size and display area of the global image according to the first pose information.
[0005] Optionally, generating a global image of the target object according to the multiple first images includes: determining second pose information of the first electronic device corresponding to the multiple first images respectively; and splicing the multiple first images according to the second pose information to form a global image of the target object.
[0006] Optionally, splicing the multiple first images according to the second pose information to form a global image of the target object includes: determining the spatial position relationship of the multiple first images according to the second pose information; and splicing the multiple first images according to the spatial position relationship of the multiple first images to form a global image of the target object.
[0007] Optionally, adjusting the size and display area of the global image according to the first pose information includes: determining the relative distance between the first electronic device and the target object according to the first pose information, and a first region image in the global image corresponding to the first pose information; using the first region image as the display area of the global image; and adjusting the size of the global image according to the relative distance.
[0008] Optionally, the video implementation method further includes: obtaining a second image sent by the first electronic device and the third pose information of the first electronic device corresponding thereto; and updating the global image according to the second image and the corresponding third pose information.
[0009] Optionally, update the global image according to the second image and the corresponding third pose information, including: determining a second region image in the global image corresponding to the third pose information; replacing the second region image with the second image to update the global image.
[0010] Optionally, the second image is a thumbnail image.
[0011] Replacing the second region image with the second image to update the global image includes: displaying the second image to the user; downloading the original image corresponding to the second image in response to a user operation; replacing the second region image with the original image corresponding to the second image to update the global image.
[0012] According to a second aspect of the present application, an embodiment of the present application provides a video implementation device, including: a first acquisition unit, configured to acquire multiple first images sent by a first electronic device, each first image being an image of a target object collected by the first electronic device, and the shooting angles of the target object in the multiple first images being different; a generation unit, configured to generate a global image of the target object according to the multiple first images; a second acquisition unit, configured to acquire first pose information sent by the first electronic device; and an adjustment unit, configured to adjust the size and the display area of the global image according to the first pose information.
[0013] According to a third aspect of the present application, an embodiment of the present application provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the video implementation method as described in the first aspect or any implementation manner of the first aspect.
[0014] According to a fourth aspect of the present application, an embodiment of the present application provides a computer-readable storage medium, the computer-readable storage medium storing computer instructions for causing a computer to execute the video implementation method as described in the first aspect or any implementation manner of the first aspect.
[0015] The video implementation method, device, electronic device, and readable storage medium provided by the embodiments of the present application obtain multiple first images sent by a first electronic device. Each first image is an image of a target object collected by the first electronic device, and the shooting angles of the target object in the multiple first images are different. A global image of the target object is generated based on the multiple first images. First pose information sent by the first electronic device is obtained. The size and display area of the global image are adjusted according to the first pose information. In this way, during a video call between the first electronic device and the second electronic device, the first electronic device does not need to send a large number of real-time images of the target object to the second electronic device. The first electronic device only needs to send multiple first images of different shooting angles of the target object to the second electronic device, and the second electronic device can generate a global image of the target object based on the multiple first images. Subsequently, during the video call process, the first electronic device only needs to synchronize the first pose information of the first electronic device to the second electronic device. Based on the first pose information, the second electronic device can determine which position of the target object the user using the first electronic device is viewing, and then adjust the size and display area of the global image based on this position, so that the user using the second electronic device can know in real time which position of the target object the user using the first electronic device is viewing. In this way, even if the network is unstable in the environment where the first electronic device is located, the same effect of the video call can be achieved.
[0016] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present application more obvious and understandable, the following specifically describes the specific embodiments of the present application. Brief Description of the Drawings
[0017] Figure 1 It is an exemplary system architecture diagram in which the embodiments of the present application can be applied;
[0018] Figure 2 It is a flowchart of a video implementation method in the embodiments of the present application;
[0019] Figure 3 It is an interaction diagram between the first electronic device and the second electronic device in the embodiments of the present application;
[0020] Figure 4 It is a schematic diagram of the first electronic device in the embodiments of the present application collecting multiple first images of a target object;
[0021] Figure 5 It is a schematic diagram of multiple first images in the embodiments of the present application;
[0022] Figure 6Schematic diagram of the display area of the global image in the embodiment of the present application;
[0023] Figure 7 Schematic diagram of the display area of the adjusted global image in the embodiment of the present application;
[0024] Figure 8 Schematic diagram of the spatial position relationship of multiple first images in the embodiment of the present application;
[0025] Figure 9 Schematic diagram of the structure of a video implementation device in the embodiment of the present application;
[0026] Figure 10 Schematic diagram of the hardware structure of an electronic device in the embodiment of the present application. Detailed implementation manners
[0027] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0028] Figure 1 An exemplary system architecture 100 to which the video implementation method or the embodiment of the video implementation device of the present application can be applied is shown.
[0029] As Figure 1 shown, the system architecture 100 may include a first electronic device 101, a network 102, and a second electronic device 103. The network 102 is used to provide a medium for a communication link between the first electronic device 101 and the second electronic device 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0030] The user may use the first electronic device 101 to interact with the second electronic device 103 through the network 102 to receive or send messages, etc. Various client applications may be installed on both the first electronic device 101 and the second electronic device 103, such as a web browser application, a shopping application, a search application, an instant messaging tool, an email client, a social platform software, etc.
[0031] The first electronic device 101 may be a wearable electronic device with a camera, a display screen, and supporting voice input, including but not limited to AR glasses.
[0032] The second electronic device 103 may be various electronic devices with a display screen and supporting voice input, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, and so on.
[0033] It should be noted that the video implementation method provided by the embodiments of the present application is generally executed by the second electronic device 103. Correspondingly, the video implementation device is generally arranged in the second electronic device 103.
[0034] Continue to refer to Figure 2 , which shows the process of a video implementation method provided by the embodiments of the present application. The video implementation method is specifically applied to the second electronic device and includes the following steps:
[0035] S201, obtain multiple first images sent by the first electronic device. Each first image is an image of a target object collected by the first electronic device, and the shooting angles of the target object in the multiple first images are different.
[0036] In this embodiment, the second electronic device is set at the far end of the first electronic device. The first electronic device communicates with the second electronic device to transmit voice and images. Among them, the user of the second electronic device, such as an expert, realizes technical guidance for the user of the first electronic device through the images displayed by the second electronic device and the voice received by the second electronic device. For example, the expert remotely guides the worker to repair, maintain, and learn static motor vehicles, etc.
[0037] The target object is the shooting object of the first electronic device. For example, it can be a motor vehicle model, an aircraft engine, etc. The first electronic device is equipped with a camera for collecting images of the target object. For a general camera, its field of view angle is not large. Therefore, the camera of the first electronic device cannot collect images of all shooting angles of the target object. Therefore, the first electronic device needs to collect multiple static images of the target object, that is, multiple first images, so that the shooting angles of the target object in the multiple first images are different.
[0038] The interaction schematic diagram between the first electronic device and the second electronic device is as Figure 3 shown.
[0039] Figure 3 In, the first electronic device collects multiple first images of the target object through its own camera.
[0040] In some embodiments, such as Figure 4 shown, the user of the first electronic device can achieve different shooting angles of the target object in the multiple first images collected by changing the standing position or the head position, and the multiple first images obtained are as Figure 5 shown. Figure 5 Among them, the multiple first images include high-definition picture 1, high-definition picture 2, and high-definition picture 3.
[0041] In some embodiments, such as Figure 5 shown, each of the first images includes the second pose information when the first electronic device collects each of the first images, that is, the 6 degrees of freedom (6DoF) position. It should be noted that how to obtain the pose information is a well-known technology that has been widely studied and applied at present, and will not be elaborated here.
[0042] Figure 3 Among them, the first electronic device sends the multiple first images collected to the second electronic device, and the second electronic device receives the multiple first images.
[0043] S202, generate a global image of the target object according to the multiple first images.
[0044] In this embodiment, the global image contains the image information when observing the target object from multiple angles. Since the multiple first images are images of multiple shooting angles of the target object, therefore, by stitching the images of multiple shooting angles, a global image of the target object can be formed. For example, if 4 first images are collected from the front, back, left, and right directions of the target object respectively, the image obtained by stitching these 4 first images is the global image. Further, image correction can also be performed on the stitching positions of each first image to obtain a panoramic image of the target object. Optionally, the global image can also refer to a three-dimensional model of the target object generated from multiple first images.
[0045] Figure 3 Among them, the second electronic device stitches the multiple first images received to generate a global image of the target object.
[0046] In some embodiments, after forming the global image of the target object, the global image can be displayed on the display screen of the second electronic device. When displaying the global image, the display area of the global image, that is, the local image of the global image, can be determined first, and then the display area is displayed. That is, the global image is a three-dimensional image, but the image displayed on the screen of the second electronic device is a planar image. When displaying the global image, the local area corresponding to the last first image sent by the first electronic device in the global image can be used as the display area of the global image, such as Figure 6 shown.
[0047] S203. Obtain the first pose information sent by the first electronic device.
[0048] In this embodiment, the first pose information is the pose information of the first electronic device itself. Optionally, the first pose information is the 6DoF information of the first electronic device.
[0049] After the first electronic device sends multiple first images collected to the second electronic device, as Figure 3 shown, the first electronic device collects its own first pose information in real time. Then, the first electronic device sends the collected first pose information to the second electronic device, and the second electronic device receives the first pose information sent by the first electronic device.
[0050] S204. Adjust the size and display area of the global image according to the first pose information.
[0051] In this embodiment, the display area of the global image is the area image of the global image displayed on the screen of the second electronic device. When the pose information of the first electronic device is different, the position of the target object seen by the user using the first electronic device is different, and the distance from the target object is different. There is a corresponding relationship between the pose information of the first electronic device and the position of the target object seen by the user, and the distance from the user to the target object. After determining the global image of the target object, based on the first pose information of the first electronic device, the position of the target object that the user of the first electronic device can see can be determined. Thus, as Figure 3 shown, the second electronic device can determine which position of the target object the user of the first electronic device is looking at based on the first pose information, and whether the user using the first electronic device is closer to or farther from the target object compared to the distance from the target object when collecting the first image. Then, the image corresponding to this position in the global image is used as the display area of the global image, and the size of the global image is adjusted correspondingly based on the analysis result of whether the user using the first electronic device is closer to or farther from the target object, so as to correspondingly adjust the size of the display area in the global image. Then, the adjusted display area of the global image is displayed on the screen for the user using the second electronic device to view, as Figure 7 shown.
[0052] Optionally, when the global image is a three-dimensional model of the target object, S204 can also be replaced with: Adjust the display angle and display size of the three-dimensional model according to the first pose information; Map the three-dimensional model into a two-dimensional image and display it according to the display angle and display size.
[0053] It should be noted that the video implementation method of the embodiments of the present application is applicable to various situations where the network state of the first electronic device is good and the network state is bad. Of course, in some embodiments, if the network state of the first electronic device is good, the first electronic device can preferentially select the normal video call method to interact with the second electronic device, that is, the first electronic device sends the dynamic image of the target object to the second electronic device in real time to achieve remote guidance. If the network state of the first electronic device is bad, the first electronic device and the second electronic device can obtain the same effect as the normal video call through the above video implementation method. Thus, when the first electronic device conducts a video with the second electronic device, it can identify the current network state in real time and convert the video mode according to the current network state.
[0054] In the video implementation method provided by the embodiments of the present application, during the video call between the first electronic device and the second electronic device, the first electronic device does not need to send a large number of real-time images of the target object to the second electronic device. The first electronic device only needs to send multiple first images of different shooting angles of the target object to the second electronic device, and the second electronic device can generate a global image of the target object based on the multiple first images. Thus, during the subsequent video call process, the first electronic device only needs to synchronize the first pose information of the first electronic device to the second electronic device. The second electronic device can determine which position of the target object the user using the first electronic device is viewing based on the first pose information, and then adjust the size and display area of the global image based on this position, so that the user using the second electronic device can know in real time which position of the target object the user using the first electronic device is viewing. In this way, even if the network in the environment where the first electronic device is located is unstable, the same effect of the video call can be achieved.
[0055] In an optional embodiment, step S202, generating a global image of the target object according to multiple first images, includes: determining the second pose information of the first electronic device corresponding to each of the multiple first images; stitching the multiple first images according to the second pose information to form a global image of the target object.
[0056] Specifically, in some implementation manners, the first electronic device can save the second pose information of the first electronic device when collecting each of the first images in each of the first images. Thus, the second electronic device can determine the second pose information of the first electronic device corresponding to each of the received images according to the received images.
[0057] In some other implementation manners, when the first electronic device sends each of the first images, it can correspondingly send the second pose information of the first electronic device corresponding to each of the images. Thus, the second electronic device obtains the second pose information of the first electronic device corresponding to each of the first images.
[0058] In an alternative embodiment, stitching a plurality of first images according to the second pose information to form a global image of the target object includes: determining the spatial position relationship of the plurality of first images according to the second pose information; stitching the plurality of first images according to the spatial position relationship of the plurality of first images to form a global image of the target object.
[0059] Specifically, the second pose information includes the position of the first electronic device in space when collecting the first image and the relative position between the target object and the first electronic device. Therefore, based on the second pose information, the position of the first image collected by the first electronic device in space can be determined. Therefore, based on the second pose information of the first electronic device corresponding to each of the plurality of images, the spatial position relationship of the plurality of first images can be determined. For example, as Figure 5 shown, each first image includes a second pose information, so that based on this second pose information, the spatial position relationship of the plurality of first images (including high-definition picture 1, high-definition picture 2, and high-definition picture 3) can be determined, as Figure 8 shown. After determining the spatial position relationship of the plurality of first images, stitching the plurality of first images according to this spatial position relationship can quickly and accurately complete the stitching of the plurality of first images and obtain a global image of the target object.
[0060] In this embodiment, by determining the second pose information of the first electronic device corresponding to each of the plurality of first images and stitching the plurality of first images according to the second pose information to form a global image of the target object, the stitching of the plurality of first images is realized by using the spatial positions of the respective first images, so that the stitching of the plurality of first images can be accurately and quickly performed to obtain a global image of the target object.
[0061] In an alternative embodiment, step S204, adjusting the size and display area of the global image according to the first pose information, includes: determining the relative distance between the first electronic device and the target object according to the first pose information, and a first region image in the global image corresponding to the first pose information; using the first region image as the display area of the global image; adjusting the size of the global image according to the relative distance.
[0062] Specifically, since each first image that makes up the global image includes the second pose information of the first electronic device, based on the first pose information of the first electronic device, the first regional image in the global image can be corresponding. By comparing the first pose information with the second pose information, it can be determined whether the relative distance between the first electronic device and the target object has increased or decreased, so as to determine whether the first regional image of the target object seen by the user of the second electronic device should be enlarged or reduced. Adjusting the size of the global image according to the relative position correspondingly adjusts the size of the first regional image, so that the adjusted first regional image is displayed on the screen of the second electronic device.
[0063] In some embodiments, after using the first regional image as the display area of the global image, the first regional image corresponding to the display area can also be intercepted, the size of the first regional image can be adjusted, and then it is displayed. Thus, the adjusted first regional image is also displayed on the screen of the second electronic device.
[0064] It should be noted that in the embodiments of the present application, the first regional image is first used as the display area of the global image, and then the size of the global image is adjusted according to the relative distance. However, in other embodiments, the execution order of these two steps can be reversed, such as first adjusting the size of the global image according to the relative distance, and then using the first regional image as the display area of the global image.
[0065] In the embodiments of the present application, the relative distance between the first electronic device and the target object is determined according to the first pose information, and the first regional image corresponding to the first pose information in the global image is used as the display area of the global image; the size of the global image is adjusted according to the relative distance, so that the specific position of the target object being viewed by the user of the first electronic device can be displayed on the second electronic device, and the first regional image corresponding to the specific position can be synchronously enlarged and reduced, and by adjusting the size of the global image, the size of the first regional image can be directly adjusted.
[0066] In an alternative embodiment, the video implementation method further includes: obtaining a second image sent by the first electronic device and the third pose information of the first electronic device corresponding thereto; updating the global image according to the second image and the corresponding third pose information.
[0067] Specifically, as Figure 3 shown, during the video communication process between the first electronic device and the second electronic device, the second image of the target object can be collected in real time at a certain frequency, for example, at a frequency of 5 images per second.
[0068] When collecting the second image of the target object, in one implementation, the user of the first electronic device can circle the key attention position in space by means of a 3D brush, so that the first electronic device only captures the local area of the target object corresponding to the key attention position subsequently. The second image is a local area map of the target object, and the size of the second image is smaller than that of the first image.
[0069] The first electronic device then compares the features of the second image with multiple first images sent to the second electronic device. If it is determined that some features in the second image have changed, the second image and the third pose information of the first electronic device corresponding to the second image can be sent to the second electronic device. Of course, before sending the second image, a prompt message can also be sent at the first electronic device end, so that the user of the first electronic device can further confirm the features that have changed in the second image. After the user confirms that it is necessary to send it to the second electronic device, the first electronic device sends the second image and the third pose information of the first electronic device corresponding to the second image to the second electronic device.
[0070] The second electronic device receives the second image sent by the first electronic device and the corresponding third pose information of the first electronic device, and updates the global image according to the second image and the corresponding third pose information.
[0071] In an alternative embodiment, updating the global image according to the second image and the corresponding third pose information includes: determining a second regional image in the global image corresponding to the third pose information; replacing the second regional image with the second image to update the global image. In this way, the global image can be updated quickly and accurately.
[0072] In the embodiments of the present application, by obtaining the second image sent by the first electronic device and the corresponding third pose information of the first electronic device, and updating the global image according to the second image and the corresponding third pose information, real-time update of the global image can be achieved.
[0073] In an alternative embodiment, in order to further reduce the bandwidth of the first electronic device, when sending the second image, the first electronic device may not send the original image of the second image, but save the original image of the second image locally and send the thumbnail image corresponding to the original image to the second electronic device. Then the second image obtained by the second electronic device is a thumbnail image.
[0074] Then replacing the second regional image with the second image to update the global image includes: displaying the second image to the user; downloading the original image corresponding to the second image in response to a user operation; replacing the second regional image with the original image corresponding to the second image to update the global image.
[0075] Specifically, a user using the second electronic device can browse the thumbnail images and select which original images corresponding to the thumbnail images to download. Thus, the second electronic device can replace the second area image with the original image corresponding to the thumbnail image to update the global image.
[0076] In the embodiment of the present application, the first electronic device sends the thumbnail images to the second electronic device. The second electronic device displays the second images to the user. In response to the user operation, the original images corresponding to the second images are downloaded, and the original images corresponding to the second images replace the second area images to update the global image. On the one hand, it can relieve the bandwidth pressure of the first electronic device. On the other hand, the user using the second electronic device can selectively download the original images corresponding to the second images, that is, can choose whether to update the global image and the update position of the global image.
[0077] The embodiment of the present application provides a video implementation device, as Figure 9 shown, including:
[0078] A first acquisition unit 21, configured to acquire multiple first images sent by the first electronic device. Each first image is an image of a target object collected by the first electronic device, and the shooting angles of the target object in the multiple first images are different.
[0079] A generation unit 22, configured to generate a global image of the target object according to the multiple first images.
[0080] A second acquisition unit 23, configured to acquire the first pose information sent by the device.
[0081] An adjustment unit 24, configured to adjust the size of the global image and the display area of the global image according to the first pose information.
[0082] In the video implementation device provided by the embodiment of the present application, during the video call between the first electronic device and the second electronic device, the first electronic device does not need to send a large number of real-time images of the target object to the second electronic device. The first electronic device only needs to send multiple first images of different shooting angles of the target object to the second electronic device. The second electronic device can generate a global image of the target object based on the multiple first images. Thus, during the subsequent video call process, the first electronic device only needs to synchronize the first pose information of the first electronic device to the second electronic device. The second electronic device can determine which position of the target object the user using the first electronic device is viewing based on the first pose information, and then adjust the size of the global image and the display area of the global image based on this position, so that the user using the second electronic device can know in real time which position of the target object the user using the first electronic device is viewing. In this way, even if the network is unstable in the environment where the first electronic device is located, the same effect of the video call can be achieved.
[0083] In some embodiments, the generating unit 22 includes:
[0084] A first determining subunit, configured to determine the second pose information of the first electronic device corresponding to multiple first images.
[0085] A stitching subunit, configured to stitch multiple first images according to the second pose information to form a global image of the target object.
[0086] In some embodiments, the stitching subunit is configured to determine the spatial position relationship of multiple first images according to the second pose information; and stitch multiple first images according to the spatial position relationship of multiple first images to form a global image of the target object.
[0087] In some embodiments, the adjusting unit 24 includes:
[0088] A second determining subunit, configured to determine the relative distance between the first electronic device and the target object according to the first pose information, and a first region image in the global image corresponding to the first pose information;
[0089] A first adjusting subunit, configured to use the first region image as the display area of the global image;
[0090] A second adjusting subunit, configured to adjust the size of the global image according to the relative distance.
[0091] In some embodiments, the video implementation device further includes:
[0092] A third obtaining unit, configured to obtain a second image sent by the first electronic device and the third pose information of the first electronic device corresponding thereto.
[0093] An updating unit, configured to update the global image according to the second image and the corresponding third pose information.
[0094] In some embodiments, the updating unit includes:
[0095] A third determining subunit, configured to determine a second region image in the global image corresponding to the third pose information.
[0096] A replacing subunit, configured to replace the second region image with the second image to update the global image.
[0097] In some embodiments, the second image is a thumbnail image. The replacing subunit is configured to display the second image to the user; in response to a user operation, download the original image corresponding to the second image; and replace the second region image with the original image corresponding to the second image to update the global image.
[0098] According to the embodiments of the present application, the present application further provides an electronic device and a readable storage medium.
[0099] Figure 10 FIG. shows a schematic block diagram of an exemplary electronic device that can be used to implement an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present application described and / or claimed herein.
[0100] As Figure 10 shown, the electronic device includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0101] Multiple components in the electronic device are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0102] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the video implementation method. For example, in some embodiments, the video implementation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the video implementation method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the video implementation method by any other suitable means (e.g., by means of firmware).
[0103] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0104] The program code for implementing the methods of this application can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0105] In the context of this application, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0106] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic, speech, or tactile input).
[0107] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0108] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server can be a cloud server, a server of a distributed system, or a server incorporating a blockchain.
[0109] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in the present application can be achieved, and no limitations are imposed herein.
[0110] In addition, the terms "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In the description of the present application, "a plurality of" means two or more, unless otherwise specifically defined.
[0111] As described above, the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.
Claims
1. A video implementation method, including: Obtain multiple first images sent by a first electronic device, where each of the first images is an image of a target object collected by the first electronic device, and the shooting angles of the target object in the multiple first images are different; Generate a global image of the target object based on the multiple first images; Obtain the first pose information sent by the first electronic device; wherein, the first pose information is the pose information of the first electronic device itself; Determine the relative distance between the first electronic device and the target object according to the first pose information, and a first region image corresponding to the first pose information in the global image; Use the first region image as the display area of the global image; Adjust the size of the global image according to the relative distance.
2. The video implementation method according to claim 1, wherein generating the global image of the target object based on the multiple first images, including: Determine the second pose information of the first electronic device corresponding to each of the multiple first images; Stitch the multiple first images according to the second pose information to form a global image of the target object.
3. The video implementation method according to claim 2, wherein stitching the multiple first images according to the second pose information to form a global image of the target object, including: Determine the spatial position relationship of the multiple first images according to the second pose information; Stitch the multiple first images according to the spatial position relationship of the multiple first images to form a global image of the target object.
4. The video implementation method according to claim 1, further including: Obtain a second image sent by the first electronic device and the third pose information of the first electronic device corresponding thereto; Update the global image according to the second image and the corresponding third pose information.
5. The video implementation method according to claim 4, wherein updating the global image according to the second image and the corresponding third pose information, including: Determine a second region image corresponding to the third pose information in the global image; Replace the second region image with the second image to update the global image.
6. The video implementation method according to claim 5, wherein the second image is a thumbnail image, and replacing the second region image with the second image to update the global image, including: Display the second image to the user; In response to a user operation, download the original image corresponding to the second image; Replace the second region image with the original image corresponding to the second image to update the global image.
7. A video implementation device, including: A first acquisition unit, configured to acquire multiple first images sent by a first electronic device, where each of the first images is an image of a target object collected by the first electronic device, and the shooting angles of the target object in the multiple first images are different; A generation unit, configured to generate a global image of the target object based on the multiple first images; A second acquisition unit, configured to acquire the first pose information sent by the first electronic device; wherein the first pose information is the pose information of the first electronic device itself; A second determination subunit, configured to determine the relative distance between the first electronic device and the target object according to the first pose information, and a first region image in the global image corresponding to the first pose information; A first adjustment subunit, configured to use the first region image as the display region of the global image; A second adjustment subunit, configured to adjust the size of the global image according to the relative distance.
8. An electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the video implementation method according to any one of claims 1-6.
9. A computer-readable storage medium storing computer instructions for causing a computer to execute the video implementation method according to any one of claims 1-6.
Citation Information
Patent Citations
Image processing method and device and storage medium
CN112073632A
Image processing method and device, electronic equipment and storage medium
CN113362227A