Image processing method and device, computer equipment and storage medium
By obtaining posture parameters on the terminal for semantic segmentation and object pose recognition, and adjusting the image to match the reference image, the image instability problem caused by terminal posture changes is solved, and the stability of the image at different sampling times is achieved.
Patent Information
- Application Number
- CN202510775577.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-26
AI Technical Summary
When the terminal captures images, the image may rotate or flip due to changes in posture, affecting image stability and resulting in a poor viewing experience for users.
By obtaining the posture parameters of the terminal at the reference moment and the first moment, semantic segmentation is performed on the reference image and the first image, image blocks are determined, and matching target image blocks are obtained from the reference image blocks. Image metadata is determined based on the object posture and adjusted to ensure image stability.
Regardless of how the terminal posture changes, image metadata adjustment can maintain the object posture matching of the target object at different sampling times, thereby improving image stability.
Smart Images

Figure CN120707849A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to an image processing method, apparatus, computer equipment, computer-readable storage medium, and computer program product. Background Art
[0002] With the rapid development of computer technology, the computing power of terminals continues to increase. Using terminals as acquisition terminals for collecting image information, due to their portability, flexibility, and ease of operation, is gaining popularity among users. To ensure consistency between shooting and playback, the acquisition terminal typically uses a fixed posture for shooting, and the playback terminal also uses a corresponding posture for playback.
[0003] However, in actual applications, the user at the acquisition end may change the shooting direction, or the acquisition end may change its posture due to poor stability. The collected images are prone to rotation or flipping, which greatly affects the user's viewing experience and has the problem of poor image stability. Summary of the Invention
[0004] Based on this, it is necessary to provide an image processing method, apparatus, computer device, computer-readable storage medium and computer program product that can improve image stability in response to the above technical problems.
[0005] In a first aspect, the present application provides an image processing method. The method comprises:
[0006] Acquire a reference image captured by a first terminal at a reference time, and a first image captured by the first terminal at a first time after the reference time;
[0007] performing semantic segmentation on the reference image and the first image based on posture parameters corresponding to the first terminal at the reference moment and the first moment, respectively, to determine a reference image block included in the reference image and a first image block included in the first image;
[0008] A target image block matching the first image block is obtained from the reference image block, and image metadata at the first moment is determined based on the object poses of the target object represented by the target image block in the reference image and the first image, respectively. The image metadata is used to adjust the first image to obtain an updated image whose object pose matches the reference image.
[0009] In a second aspect, the present application also provides another image processing method. The method comprises:
[0010] Obtaining image data corresponding to a plurality of consecutive sampling moments;
[0011] For each of the sampling moments, when the image data corresponding to the sampling moment includes image metadata and image information, the image information is adjusted based on the image metadata to obtain an updated image corresponding to the sampling moment; the image metadata is obtained based on the above-mentioned image processing method.
[0012] In a third aspect, the present application further provides an image processing device. The device comprises:
[0013] an image acquisition module, configured to acquire a reference image acquired by a first terminal at a reference time, and a first image acquired by the first terminal at a first time after the reference time;
[0014] a semantic segmentation module, configured to perform semantic segmentation on the reference image and the first image based on the posture parameters corresponding to the first terminal at the reference moment and the first moment, respectively, to determine a reference image block included in the reference image and a first image block included in the first image;
[0015] a metadata determination module configured to obtain, from the reference image block, a target image block that matches the first image block, and determine image metadata at the first moment based on the object poses of a target object represented by the target image block in the reference image and the first image, respectively; the image metadata being used to adjust the first image to obtain an updated image whose object pose matches the reference image.
[0016] In a fourth aspect, the present application further provides another image processing device. The device comprises:
[0017] An image data acquisition module is used to acquire image data corresponding to a plurality of consecutive sampling moments;
[0018] An image adjustment module is configured to adjust the image information based on the image metadata for each sampling moment, when the image data corresponding to the sampling moment includes image metadata and image information, to obtain an updated image corresponding to the sampling moment; the image metadata is obtained based on the above-mentioned image processing method.
[0019] In a fifth aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned image processing method when executing the computer program.
[0020] In a sixth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned image processing method when executed by a processor.
[0021] In a seventh aspect, the present application further provides a computer program product, which includes a computer program that implements the above-mentioned image processing method when executed by a processor.
[0022] The above-described image processing method, apparatus, computer device, computer-readable storage medium, and computer program product obtain a reference image captured by a first terminal at a reference time, and a first image captured by the first terminal at a first time after the reference time; perform semantic segmentation on the reference image and the first image based on the posture parameters corresponding to the first terminal at the reference time and the first time, respectively, to determine the reference image blocks contained in the reference image and the first image blocks contained in the first image; obtain a target image block matching the first image block from the reference image block, and determine image metadata for the first time based on the object poses of the target object represented by the target image block in the reference image and the first image, respectively. In the above process, image metadata used to adjust the first image is determined through semantic segmentation and object pose recognition, which can ensure the accuracy of the image metadata. Therefore, regardless of how the posture parameters of the first terminal change, the first image captured at the first time can be adjusted based on the image metadata corresponding to the first time, obtaining an updated image whose object pose matches the reference image. This ensures that the object poses of the same target object at different sampling times can match, which is beneficial for improving image stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A diagram showing an application environment of an image processing method in one embodiment;
[0024] Figure 2 1 is a flow chart of an image processing method according to an embodiment;
[0025] Figure 3 Schematic diagram of the attitude angle of a mobile phone in one embodiment;
[0026] Figure 4 Schematic diagram of image semantic segmentation results in one embodiment;
[0027] Figure 5 is a schematic diagram of an image processing process in one embodiment;
[0028] Figure 6 Schematic diagram of object posture change caused by posture angle change in one embodiment;
[0029] Figure 7 A schematic diagram of priorities configured for different object categories in one embodiment;
[0030] Figure 8 is a schematic diagram of an image adjustment process in one embodiment;
[0031] Figure 9 is a schematic diagram of image metadata in one embodiment;
[0032] Figure 10 is a schematic diagram of image metadata in another embodiment;
[0033] Figure 11 is a flowchart of an image processing method in another embodiment;
[0034] Figure 12 is a schematic diagram of an image adjustment process in another embodiment;
[0035] Figure 13 is a flowchart of an image processing method in yet another embodiment;
[0036] Figure 14 is a structural block diagram of an image processing device in one embodiment;
[0037] Figure 15 is a structural block diagram of an image processing device in another embodiment;
[0038] Figure 16 is a diagram of the internal structure of a computer device in one embodiment;
[0039] Figure 17 FIG. 4 is a diagram showing the internal structure of a computer device in another embodiment. DETAILED DESCRIPTION
[0040] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0041] The image processing method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the first terminal 110 and the second terminal 120 can communicate with the server 130 through a network. The communication network can be a wired network or a wireless network. Therefore, the first terminal 110 and the server 130 can be directly or indirectly connected through wired or wireless communication. For example, the first terminal 110 can be indirectly connected to the server 130 through a wireless access point, or the first terminal 110 can be directly connected to the server 130 through the Internet, and this application does not limit this.
[0042] The first terminal 110 and the second terminal 120 are terminal devices with image acquisition capabilities, and may include, but are not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices may include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices may include smart watches, smart bracelets, and head-mounted devices. The server 130 may be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The data storage system may store data that the server 130 needs to process. The data storage system may be separate, integrated with the server 130, or located in the cloud or on other devices.
[0043] It should be noted that the image processing method in the embodiment of the present application can be executed by the first terminal 110, the second terminal 120 or the server 130 individually, or can be executed together through multi-terminal interaction. Figure 1 As shown, the first terminal 110 may refer to an image acquisition terminal, which is used to continuously acquire image information at different sampling times; the second terminal 120 may refer to an image playback terminal, which is used to obtain and play the image information acquired by the first terminal 110. The image processing method provided in this application can be applied to scenes with relatively fixed shooting content, such as live broadcasts, news broadcasts, video conferences, speeches, and interviews. Among them, during live broadcasts and news broadcasts, the shooting content usually includes the anchor; the shooting content of video conferences usually includes the meeting participants; the shooting content of speeches usually includes the speaker; and the shooting content during interviews usually includes the interviewee.
[0044] Take the case where the server 130 executes alone as an example. In an optional embodiment, the server 130 can obtain the reference image 111 captured by the first terminal 110 at the reference moment, the first image 112 captured by the first terminal 110 at the first moment after the reference moment, and the posture parameters corresponding to the first terminal 110 at the reference moment and the first moment respectively. Then, based on the posture parameters corresponding to the first terminal 110 at the reference moment and the first moment respectively, semantic segmentation is performed on the reference image 111 and the first image 112 respectively to determine the reference image block contained in the reference image 111 and the first image block contained in the first image 112. Next, the server 130 can perform matching analysis on the reference image block and the first image block, obtain the target image block that matches the first image block from the reference image block, and determine the image metadata at the first moment based on the object posture corresponding to the target object represented by the target image block in the reference image 111 and the first image 112 respectively. Exemplarily, the target object can be, for example, Figure 1 The image metadata may include an image rotation angle determined based on the change in the posture of object A. Thus, the server 130 can adjust the first image 112 based on the image metadata, obtain an updated image 122 whose object posture matches the reference image 111, and send it to the second terminal 120. This ensures that the object posture of the target object in the content played on the second terminal 120 is relatively stable.
[0045] In an optional embodiment, if the performance of the terminal meets the image processing requirements, the first terminal 110 can also calculate the image metadata and adjust the image based on the image metadata, so that the second terminal 120 can obtain and play image information at different times through the server 130.
[0046] In an optional embodiment, the second terminal 120 can also obtain the image information collected by the first terminal 110 and the posture parameters of the first terminal 110 through the server 130, and perform semantic segmentation and matching on the image information based on the posture parameters, determine the image metadata, and adjust the image based on the image metadata, and finally play the updated image after adjustment.
[0047] In an optional embodiment, the first terminal 110 collects image information and reports posture parameters corresponding to the image information to the server 130. The server 130 performs semantic segmentation and matching on the image information based on the posture parameters, determines image metadata, and sends the image data corresponding to the first image, including the image metadata, to the second terminal 120. The second terminal 120 obtains image data corresponding to multiple consecutive sampling moments. For each sampling moment, if the image data corresponding to the sampling moment includes image metadata and image information, the image information is adjusted based on the image metadata, and an updated image corresponding to the sampling moment is obtained and played.
[0048] In one embodiment, an image processing method is provided, which can be applied to a computer device, which can be a terminal or a server. Figure 1 The server 130 in FIG. 1 is taken as an example to illustrate, in this embodiment, Figure 2 As shown, the method includes the following steps:
[0049] Step S202: Acquire a reference image captured by the first terminal at a reference time and a first image captured by the first terminal at a first time after the reference time.
[0050] The reference moment and the first moment are both sampling moments. The reference moment is a moment used as a benchmark or starting point for other operations. The reference image is the image captured by the first terminal at the reference moment, which provides a reference or benchmark. The first moment is the sampling moment after the reference moment. This first moment can be the sampling moment immediately following the reference moment, or there can be other sampling moments between the first moment and the reference moment. The first image is the image captured by the first terminal at the first moment.
[0051] Specifically, the first terminal has image acquisition capabilities and can acquire corresponding image information at multiple sampling moments. Optionally, the first terminal can periodically acquire images according to an acquisition time interval, which can be, for example, 300ms or 500ms. The server can obtain the image information acquired by the first terminal at the multiple sampling moments and determine a reference moment from each sampling moment to obtain a reference image acquired by the first terminal at the reference moment and a first image acquired by the first terminal at the first moment after the reference moment.
[0052] Optionally, the first terminal may upload the collected image information to the server, so that the server may passively receive the image information; or the server may actively obtain the image information from the first terminal.
[0053] In one possible implementation, the first terminal may be a live broadcast terminal, and the image information collected by the first terminal may be sent to the server via a live stream. Optionally, the server may perform a delay process on the received live stream and, during the delay process, employ the image processing method provided in this application to perform image processing on the acquired image information, thereby ensuring the smoothness of the live stream sent to the second terminal.
[0054] Optionally, the reference time can be specified by the user of the first terminal to represent the sampling time under the correct sampling posture recognized by the user. The correct sampling posture can be, for example, the posture of the device when the first terminal is placed vertically. For example, if the first terminal is a mobile phone, Figure 3 As shown, the posture of the mobile phone can be represented by the posture angle, and specifically by three Euler angles to represent the rotation state of the mobile phone around three different axes. Figure 3 As shown, the three axes can correspond to the long side (X-axis), short side (Y-axis) and vertical direction (Z-axis) of the mobile phone respectively, and the intersection of these three axes is the center of the mobile phone camera 310. Furthermore, the posture angle can include pitch, yaw and roll. Among them, the pitch angle represents the rotation angle of the mobile phone around its roll axis (X-axis); the yaw angle represents the rotation angle of the mobile phone around its vertical axis (Z-axis); and the roll angle represents the rotation angle of the mobile phone around its pitch axis (Y-axis). The correct sampling posture can refer to vertical screen shooting, that is, the mobile phone is placed vertically and all Euler angles are 0°.
[0055] Optionally, since the sampling posture at the first sampling moment is usually the posture of the user of the first terminal after adjustment, the server may use the first sampling moment as a reference moment and the image captured at the first sampling moment as a reference image.
[0056] Step S204 : Based on the posture parameters corresponding to the first terminal at the reference moment and the first moment, semantic segmentation is performed on the reference image and the first image to determine the reference image block included in the reference image and the first image block included in the first image.
[0057] The posture parameters may include posture angles, translation parameters, and a posture matrix. The posture angle is used to describe the rotational state of the first terminal in three-dimensional space; the translation parameter is used to describe the translational position of the first terminal in three-dimensional space; and the posture matrix may include translation and rotation parameters, which can be used to describe the rotational state and translational position of the first terminal in three-dimensional space. Optionally, if the first terminal is a mobile phone, the posture parameters may include the phone's Euler angles; if the first terminal is a six-degree-of-freedom camera, the posture parameters may include translation parameters and posture angles, which are used to record the posture information of the six-degree-of-freedom camera in six degrees of freedom (corresponding to three translation directions and three rotation directions).
[0058] Image semantic segmentation is a core task in computer vision. Its goal is to assign each pixel in an image to a predefined category label, thereby achieving pixel-level understanding of the image content. Unlike traditional image segmentation methods, semantic segmentation focuses not only on demarcating image boundaries but also emphasizes the semantic classification of objects or scenes in the image, that is, identifying the category to which each pixel belongs. For example, in an image containing a news anchor and news subtitles, the anchor's pixels can be labeled "person," the subtitles' pixels can be labeled "text," and the background can be labeled "background."
[0059] The image block is the segmentation result of semantic segmentation. Optionally, the image is segmented into one or more image blocks by semantic segmentation, wherein pixels belonging to the same category and adjacent to each other are divided into the same image block, and the image before segmentation can be obtained by splicing the image blocks. Figure 4 As shown, semantic segmentation is performed on image 410 to obtain image block 411, image block 412 and image block 413, wherein the category corresponding to image block 411 is "person", the category corresponding to image block 412 is "text", and the category corresponding to image block 413 is "background".
[0060] Specifically, the server can perform semantic segmentation on the reference image and the first image based on the posture parameters corresponding to the first terminal at the reference moment and the first moment, respectively, to determine the reference image block contained in the reference image and the first image block contained in the first image. Optionally, the posture parameters can be used as a trigger condition for semantic segmentation, that is, the server does not need to perform image semantic segmentation processing for each sampling moment, but performs image semantic segmentation processing when it is determined that the trigger condition for semantic segmentation is met according to the posture parameters, and determines the image metadata based on the semantic segmentation result. Optionally, the posture parameters can also be used as parameters in the semantic segmentation process, that is, for each sampling moment after the reference moment, the sampling moment can be used as the first moment, and semantic segmentation can be performed on the image collected at the sampling moment to determine the image metadata corresponding to the sampling moment.
[0061] In an optional embodiment, the server can determine whether the first moment meets the trigger condition of semantic segmentation based on the posture parameters corresponding to the first terminal at the reference moment and the first moment respectively, and perform semantic segmentation on the reference image and the first image respectively when the trigger condition is met at the first moment. Exemplarily, the server can determine the sampling posture change of the first terminal at the first moment relative to the reference moment, and perform semantic segmentation processing when the sampling posture change meets the change condition. Since the sampling posture change is the direct cause of image stability, performing semantic segmentation on images with large sampling posture changes, and determining image metadata for guiding image adjustment based on the segmentation results can improve image stability while reducing the consumption of computing resources.
[0062] In an optional embodiment, the server may determine the depth information of each pixel in the reference image and the first image based on the posture parameters of the first terminal at the reference time and the first time, respectively, and further perform semantic segmentation on the reference image and the first image based on the color information and depth information of each pixel to improve the accuracy of the semantic segmentation results. The color information may be, for example, color information in a color space such as RGB (Red Green Blue) or HSV (Hue Saturation Value).
[0063] Taking the semantic segmentation process of a reference image as an example, the server can optionally perform feature extraction on the reference image's color information to obtain color features, and feature extraction on the reference image's depth information to obtain depth features. Then, by fusing the color and depth features, a fused feature of the reference image is obtained, and the semantic segmentation model is used to perform semantic segmentation on the reference image using this fused feature. Specific methods for fusing color and depth features include concatenation, weighted fusion, or an attention mechanism. Semantic segmentation models can include fully convolutional neural networks (FCNs), DeepLab networks, pyramid scene parsing networks (PSPNets), or U-Nets. Optionally, after extracting the reference image's color features, the server can use depth information as weights to enhance or suppress features in certain areas to improve the accuracy of the semantic segmentation results.
[0064] In an optional embodiment, the server can also determine whether the first moment meets the trigger conditions for semantic segmentation based on the posture parameters corresponding to the first terminal at the reference moment and the first moment respectively, and if the trigger conditions are met at the first moment, further perform semantic segmentation on the reference image and the first image based on each posture parameter, so as to improve the accuracy of the semantic segmentation results while saving computing resources.
[0065] It is understandable that the server can use the same or different semantic segmentation methods for the reference image and the first image. Optionally, the server can use the same semantic segmentation method to perform semantic segmentation on the reference image and the first image respectively to ensure the comparability of the segmentation results. Optionally, the server can detect edges in the image based on an edge detection algorithm and connect the edges to form segmented image blocks; the server can also divide the image into multiple image blocks based on pixel similarity; the server can also perform semantic segmentation on the image based on deep learning methods to obtain multiple image blocks.
[0066] Step S206 , obtaining a target image block matching the first image block from the reference image block, and determining image metadata at the first moment based on the corresponding object poses of the target object represented by the target image block in the reference image and the first image, respectively.
[0067] The image metadata is used to adjust the first image to obtain an updated image in which the object pose matches the reference image. The image metadata may include core image information, image rotation angle, etc. The core image information may include the object position of the object to be highlighted in the image, as well as the image size, etc. Optionally, the object to be highlighted in the image may be, for example, a person, text, etc. The image size may include, for example, width and height. The image rotation angle may refer to the rotation angle at the first moment relative to the reference moment, or it may refer to the absolute rotation angle at the first moment relative to the rotation axis.
[0068] The object pose matching between the updated image and the reference image means that the object poses of the target object in the reference image and the updated image meet the pose similarity condition. Figure 1 As shown, the position of the object A in the reference image 111 and the updated image 122 is close to the center of the page, and the posture of the object A in the reference image 111 and the updated image 122 is close to a vertical state.
[0069] Specifically, the server may extract image features of the reference image block and the first image block respectively, perform image feature matching on the reference image block and the first image block, and obtain a target image block that matches the first image block from the reference image block.
[0070] In an optional embodiment, the server may perform feature matching on reference image blocks of a predetermined category and the first image block based on the categories of the image blocks, and obtain a target image block from the reference image blocks that matches the first image block. The predetermined category may be set by the user of the first terminal and is used to identify the category of an object that is desired to be highlighted during the capture process. For example, the predetermined category may be "people" or "text."
[0071] In an optional embodiment, each image block obtained by semantic segmentation is an image block of the foreground portion. In this embodiment, the server can extract the image features of each image block and calculate feature similarity for the image features of the reference image block and the first image block of the same category, thereby determining the image block combination with the highest similarity and determining the reference image block in this image block combination as the target image block.
[0072] It is understandable that in scenarios with relatively fixed content, such as live broadcasts, news broadcasts, video conferences, speeches, and interviews, the objects that appear in images captured at different sampling times are generally the objects that the first terminal desires to highlight during the shooting process. Based on this, once the target image block has been determined, the server can further determine the target object represented by the target image block and, based on the corresponding object poses of the target object in the reference image and the first image, determine the image metadata used for image adjustment of the first image.
[0073] Optionally, the server may determine the feature position at the first moment according to the object position of the target object in the first image, and obtain image metadata including the feature position.
[0074] Optionally, the server can determine the image rotation angle at the first moment relative to the reference moment based on the object postures of the target object in the reference image and the first image, respectively, and obtain image metadata containing the image rotation angle. Optionally, the image rotation angle can belong to the angle set {90, -90, 180, -180}. In this case, the server can determine the object rotation angle of the target object at the first moment relative to the reference moment based on the posture change of the target object in the reference image relative to the first image, and further select the image rotation angle whose opposite number is closest to the object rotation angle from the angle set. For example, Figure 1 As shown, the rotation angle of object A at the first moment relative to the reference moment is 80°. The corresponding image rotation angle at the first moment is -90°. The opposite of this image rotation angle, 90°, is closest to the object rotation angle of 80°. Since the photosensitive chip on the acquisition end and the screen on the playback end are typically rectangular, selecting an image rotation angle from the angle set eliminates the need to crop the rotated image, improving efficiency during the image update process.
[0075] Optionally, if there is no matching reference image block and first image block, the server may not need to calculate the image metadata for the first moment, but instead take the next sampling moment of the first moment as the new first moment, and perform image processing on the new first moment based on the reference moment.
[0076] The image processing method described above obtains a reference image captured by a first terminal at a reference moment and a first image captured by the first terminal at a first moment after the reference moment; performs semantic segmentation on the reference image and the first image based on the posture parameters corresponding to the first terminal at the reference moment and the first moment, respectively, to determine the reference image blocks contained in the reference image and the first image blocks contained in the first image; obtains a target image block matching the first image block from the reference image block, and determines image metadata for the first moment based on the object poses of the target object represented by the target image block in the reference image and the first image, respectively. In the above process, image metadata used to adjust the first image is determined through semantic segmentation and object pose recognition, which can ensure the accuracy of the image metadata. Therefore, regardless of how the posture parameters of the first terminal change, the first image captured at the first moment can be adjusted based on the image metadata corresponding to the first moment, obtaining an updated image whose object pose matches the reference image. This ensures that the object poses of the same target object at different sampling moments can match, which is beneficial for improving image stability.
[0077] In one embodiment, semantic segmentation is performed on the reference image and the first image based on the posture parameters corresponding to the first terminal at the reference moment and the first moment, respectively, including: obtaining the posture parameters corresponding to the first terminal at the reference moment and the first moment, respectively; determining the sampling posture change of the first terminal at the first moment relative to the reference moment based on each posture parameter; and when the sampling posture change meets the change condition, semantic segmentation is performed on the reference image and the first image, respectively.
[0078] The posture parameters may include posture angles, translation parameters, and posture matrices. The change in the sampled posture may be characterized by a change in the value of at least one posture parameter. The change in the sampled posture meeting the change condition may mean that the change in the sampled posture is greater than or equal to a change threshold.
[0079] Specifically, the server can obtain the posture parameters corresponding to the first terminal at the reference time and the first time. Figure 5As shown, in response to user operations, the first terminal can open its own sensor data permissions and encode the sensor data into the corresponding SEI (Supplemental Enhancement Information). This allows data integration to obtain image data containing the captured image information and the SEI information corresponding to the image information. The server can then analyze the image data to obtain the posture parameters corresponding to the first terminal at the sampling time while simultaneously obtaining the image information captured at the sampling time. Optionally, the sensor that collects the sensor data may include at least one of a gyroscope, a magnetometer, an orientation sensor, or an accelerometer.
[0080] Take the case where the first terminal is the terminal and the posture parameters include Euler angles as an example. Figure 2 and Figure 6 As shown, a change in the pitch angle of the first terminal will cause the position of object A in the image to shift left or right; a change in the roll angle of the first terminal will cause the position of object A in the image to shift up or down; and a change in the yaw angle of the first terminal will cause object A to rotate in the image. In either case, the posture of object A will change, thereby affecting image stability. Based on this, when the posture parameters are obtained, the server can determine the change in the sampling posture of the first terminal at the first moment relative to the reference moment based on each posture parameter, and perform semantic segmentation on the reference image and the first image respectively if the sampling posture change meets the change condition.
[0081] Optionally, the server can determine the parameter value change of the first terminal at the first moment relative to the reference moment for each posture parameter, obtain the sampling posture change represented by the parameter value change corresponding to the multiple posture parameters, and determine that the sampling posture change meets the change condition when at least one parameter value change meets the parameter value change condition. The parameter value change conditions corresponding to different posture parameters can be the same or different. Exemplarily, when the first terminal is a mobile phone, the sampling posture change can be represented by the pitch angle change, the yaw angle change, and the roll angle change, and when any one of the angle changes is greater than or equal to the corresponding angle change threshold, it is determined that the sampling posture change of the mobile phone meets the change condition. The angle change threshold can be, for example, 30°, 40°, or 45°.
[0082] In a possible implementation, considering that object rotation has a greater impact on the user's viewing experience, while the change of object position has a smaller impact on the user's viewing experience, the server can also Figure 2If the yaw angle change is greater than or equal to the corresponding angle change threshold, the phone's sampled posture change is determined to meet the change condition. In this case, the image metadata determined by image semantic segmentation is used to rotate the first image to obtain an updated image in which the object's posture matches the reference image.
[0083] Optionally, the server can also obtain a statistical change by counting the parameter value changes corresponding to multiple posture parameters, and obtain the sampling posture change represented by the statistical change. The way of counting the parameter value changes can be averaging, summing, etc., which is not limited here. For example, when the first terminal is a mobile phone, the statistical change can be obtained by counting the pitch angle change, yaw angle change and roll angle change, and when the statistical change is greater than or equal to the corresponding statistical change threshold, it is determined that the sampling posture change of the mobile phone meets the change condition. The statistical change threshold can be, for example, 45° or 50°.
[0084] It can be understood that when the sampling posture change does not meet the change condition, the server does not need to perform semantic segmentation. At this time, the change of the first image relative to the reference image is relatively small, the image stability is relatively good, and the impact on the user's viewing experience is relatively small.
[0085] Optional, such as Figure 5 As shown, the server obtains image data uploaded by the first terminal and analyzes the sensor data in the image data to determine whether the change in the sampling posture of the first terminal at the first moment relative to the reference moment meets the change condition. If the sampling posture change meets the change condition, the image understanding function is activated, semantic segmentation and feature matching are performed, image metadata at the first moment is determined, and the image data containing the image metadata is sent to the second terminal. If the sampling posture change does not meet the change condition, the image data without the image metadata is directly sent to the second terminal. Taking the case where the reference moment is the first sampling moment as an example, the server can use the second sampling moment as the first moment in chronological order. If the image understanding function is not activated, or the reference image block does not match the first image block, the first image can be directly sent to the second terminal, and the third sampling moment can be used as the new first moment. If the server determines the image metadata at the second sampling moment, it can combine the image metadata with the first image into image data and send it to the second terminal. The third sampling moment is used as the new first moment for the next round of image processing.
[0086] Optional, such as Figure 5 As shown, in the case where the image data includes image metadata, the second terminal can analyze the image metadata in the image data, adjust the first image captured at the first moment, obtain an updated image, and play it.
[0087] In the above embodiment, when the sampling posture change meets the change condition, semantic segmentation is performed on the reference image and the first image respectively. Since the sampling posture change is the direct cause of affecting the image stability, semantic segmentation is performed on the image with a large sampling posture change, and the image metadata used to guide image adjustment is determined based on the segmentation results. This can improve image stability while reducing the consumption of computing resources.
[0088] In an optional embodiment, the attitude parameter includes an attitude angle. In this embodiment, obtaining the attitude parameters corresponding to the first terminal at the reference time and the first time, respectively, includes: obtaining gyroscope information corresponding to the first terminal at the reference time and the first time, respectively; and determining the attitude angle of the first terminal based on the gyroscope information.
[0089] The gyroscope information is data collected by a gyroscope built into the first terminal. A gyroscope is an instrument used to measure and maintain directional stability. Its core principle is based on the gyroscopic effect, which determines the direction of an object by measuring the angular velocity of rotation. Specifically, in this application, the gyroscope information is used to characterize the device orientation of the first terminal. The attitude angle of the first terminal may include at least one of a pitch angle, a roll angle, or a yaw angle.
[0090] Optionally, the server may obtain gyroscope information corresponding to the first terminal at the reference moment and the first moment, respectively, and extract the attitude angle of the first terminal from the gyroscope information.
[0091] Optionally, when the gyroscope information is a quaternion, the server may calculate the attitude angle of the first terminal based on the quaternion.
[0092] Taking the quaternion Q=[q0,q1,q2,q3] as an example, the corresponding attitude angle can be expressed as:
[0093]
[0094] Among them, q0, q1, q2, q3 are the four components of the quaternion, pitch is the pitch angle, roll is the roll angle, and yaw is the yaw angle.
[0095] In the above embodiment, the attitude angle of the first terminal is determined based on the gyroscope information. Thanks to the advantages of the gyroscope such as high dynamic response, high precision, strong real-time performance, and applicability to complex environments, the accuracy and timeliness of the attitude angle can be improved.
[0096] In practical applications, to ensure that the target object to be matched is the object that is expected to be highlighted during the image acquisition process, the server can perform semantic segmentation on the image to obtain the reference image block and the first image block belonging to the foreground image. Specifically, the server can first perform semantic segmentation on the image, then identify the image blocks belonging to the foreground image from the segmentation results to obtain the reference image block and the first image block belonging to the foreground image. Alternatively, the server can first identify the foreground image and then perform semantic segmentation on the foreground image to obtain the reference image block and the first image block belonging to the foreground image.
[0097] In an optional embodiment, semantic segmentation is performed on the reference image and the first image respectively to determine the reference image blocks contained in the reference image and the first image blocks contained in the first image, including: performing semantic segmentation on the reference image and the first image respectively to determine multiple reference candidate image blocks contained in the reference image and multiple first candidate image blocks contained in the first image; determining the reference image blocks belonging to the foreground image from each reference candidate image block; and determining the first image blocks belonging to the foreground image from each first candidate image block.
[0098] The foreground image refers to the portion of the image that is at the forefront, closest to the viewer, and visually the most prominent and eye-catching. It is typically the primary object or element in the image, carrying the key information the image is intended to convey. The background image refers to the portion of the image that lies behind the foreground, further from the viewer, and provides context and atmosphere for the foreground. It is typically a secondary element in the image, serving to highlight and enhance the foreground image.
[0099] Specifically, the server can first perform semantic segmentation on the reference image and the first image respectively, and determine the multiple reference candidate image blocks contained in the reference image, and the multiple first candidate image blocks contained in the first image. Then, from each reference candidate image block, identify the reference image block belonging to the foreground image, and from each first candidate image block, identify the first image block belonging to the foreground image. Optionally, the server can remove the image blocks belonging to the background from each candidate image block based on the category to which each candidate image block belongs, and obtain the reference image block and the first image block. Optionally, according to a predefined foreground category, the reference image blocks and the first image blocks belonging to the foreground category can be filtered from each candidate image block. The foreground category may include, for example, "person", "text", etc. Exemplary, such as Figure 4 As shown, the identified candidate image blocks may include image block 411 , image block 412 , and image block 413 , and the image blocks belonging to the foreground image include image block 411 and image block 412 .
[0100] In an optional embodiment, semantic segmentation is performed on the reference image and the first image respectively to determine the reference image blocks contained in the reference image and the first image blocks contained in the first image, including: performing foreground image recognition on the reference image and the first image respectively to determine the reference foreground image corresponding to the reference image and the first foreground image corresponding to the first image; and semantic segmentation is performed on the reference foreground image and the first foreground image respectively to determine the reference image blocks contained in the reference image and the first image blocks contained in the first image.
[0101] Specifically, the server can convert the image into a color space such as HSV or RGB and set a color threshold to identify the foreground area in the image, thereby obtaining a reference foreground image corresponding to the reference image and a first foreground image corresponding to the first image. The server can also implement foreground image recognition based on a deep learning model by learning the feature differences between the foreground image and the background image. After completing foreground image recognition, the server can perform semantic segmentation on the reference foreground image to determine the reference image blocks contained in the reference image, and perform semantic segmentation on the first foreground image to determine the first image blocks contained in the first image.
[0102] In the above embodiment, semantic segmentation is used to obtain a reference image block and a first image block belonging to the foreground image, and image matching is performed on this basis. This can ensure that the matched target object is the object that is desired to be highlighted during the image acquisition process, thereby ensuring the accuracy of the image metadata determined based on the object posture, which is conducive to further improving image stability and enhancing the user experience on the playback end.
[0103] In an optional embodiment, there are multiple reference image blocks and multiple first image blocks. In this embodiment, the image processing method further includes: determining the priority of each first image block based on the object category of the object depicted by each first image block; determining a selected image block with the highest priority from each first image block, and performing image feature matching on the selected image block with each reference image block; if a matching image block exists in each reference image block that matches the selected image block, determining the first image block as the target image block; if none of the reference image blocks matches the selected image block, determining the selected image block with the highest priority from the remaining first image blocks, and returning to the step of performing image feature matching on the selected image block with each reference image block.
[0104] Among them, the object category is determined by semantic segmentation, which can include coarse-grained categories such as "foreground" and "background", as well as fine-grained categories such as "people", "animals", "objects", and "text". Figure 7As shown, you can configure corresponding priorities for different categories, where "foreground" has a higher priority than "background," and "people" in the foreground have the highest priority. Optionally, this priority represents the importance of the object category and can be configured by the user of the first terminal or by the server based on expert experience.
[0105] Specifically, the server can determine the priority of each first image block based on the object category of the object depicted by each first image block, and then determine the selected image block with the highest priority from each first image block, and obtain the image features of the selected image block and each reference image block through image feature extraction. Then, based on each image feature, the selected image block is matched with each reference image block. If there is an image block that matches the selected image block in each reference image block, then the image block is determined as the target image block. If none of the reference image blocks matches the selected image block, then the selected image block with the highest priority is determined from the remaining first image blocks, and the process of matching the selected image block with each reference image block is returned to, until the target image block is obtained through screening, or until all the first image blocks have been traversed.
[0106] For example, Figure 4 As shown, if image 410 is the first image, image block 411 has the highest priority. The server can first perform image feature matching on image block 411 with reference image blocks. If a matching image block exists among the reference image blocks, that image block is determined as the target image block. If none of the reference image blocks matches image block 411, image block 412 is then matched with the reference image blocks for image feature matching until the target image block is found or all image blocks in image 410 have been traversed.
[0107] If the target image block is still not obtained after all the first image blocks have been traversed, it means that the first image and the reference image do not contain the same object, and the image metadata cannot be determined through object pose matching. In this case, the server can determine the image metadata at the first moment based on the change in the sampling posture of the first terminal at the first moment relative to the reference moment. For example, the server can determine the inverse of the change in the yaw angle of the first terminal at the first moment relative to the reference moment as the image rotation angle at the first moment. It can be understood that the accuracy of the image metadata determined based on object pose matching is better than the accuracy of the image metadata determined based on the change in sampling posture. Therefore, when the image metadata cannot be determined through object pose matching, the server may not calculate the image metadata for the first moment, but instead use the next sampling moment after the first moment as the new first moment and perform image processing on the new first moment based on the reference moment.
[0108] It's important to note that there's no single way to extract image features from image blocks. For example, it can be implemented using deep learning methods like convolutional neural networks and region-based convolutional neural networks, or traditional methods like gray-level co-occurrence matrices and color histograms. Alternatively, MobileNet can be used for image feature extraction. MobileNet is an efficient convolutional neural network architecture designed specifically for image recognition tasks running on mobile devices and embedded systems. It uses depthwise separable convolution to reduce computational effort and the number of parameters, achieving lightweight performance while maintaining high accuracy.
[0109] In the above embodiment, image feature matching is performed in order of priority, and a target image block matching the first image block is obtained from the reference image blocks. It is not necessary to perform image feature matching on all image blocks, which can save computing resources and improve work efficiency.
[0110] In one embodiment, based on the object postures of the target object represented by the target image block in the reference image and the first image, respectively, determining the image metadata at the first moment includes: determining the object position of the target object represented by the target image block in the first image as the feature position at the first moment; based on the object postures of the target object in the reference image and the first image, respectively, determining the image rotation angle at the first moment relative to the reference moment; and determining the image metadata including the feature position and the image rotation angle.
[0111] Among them, the target object is an object that exists in both the reference image and the first image, and the target object is usually also the object that is expected to be highlighted during the image acquisition process. Based on this, the server can determine the target object represented by the target image block, and then determine the object position of the target object in the first image, and determine the object position as the feature position at the first moment. It can be understood that the feature position is the position where the object is expected to be highlighted in the first image. Therefore, during the image adjustment process, the feature position can be translated to the picture feature position to ensure the image display effect. The picture feature position can be, for example, the center of the picture. Optionally, in the process of image adjustment based on the feature position, the image block to which the feature position belongs can be moved, or the entire image can be translated.
[0112] Optionally, when translating the entire image, edge images may be lost. Therefore, the server can zoom in and crop the translated image to ensure image integrity. In one possible implementation, the zoom factor during the zoom process can be determined based on the ratio of the distance between the feature position and the image edge closest to the feature position to the distance between the desired position after translation and the image edge. Figure 8 As shown, it is possible to Figure 8 After translating the first image 801 by the object position of the object A in the image, the translated image 802 is enlarged and cropped to obtain an updated image 803 that is consistent with the size of the original image. Figure 8 In the example, the magnification factor can be expressed as scale = d2 / d1, where d1 represents the distance between the feature position and the right edge of the image closest to the feature position, and d2 represents the distance between the desired position after translation and the right edge of the image. In this case, after translating to the left and magnifying the image, the top, left, and bottom edges of the image are all cropped.
[0113] On the other hand, the server can also determine the object rotation angle of the target object at the reference moment relative to the first moment based on the object postures of the target object in the reference image and the first image, respectively, and further determine the image rotation angle used to compensate for the object rotation angle. Optionally, the object position may refer to the position of the centroid of the target object in the first image, or may refer to the center position of the first image block where the target object is located. Optionally, the object position may be represented by coordinates. Optionally, the image rotation angle may be the reciprocal of the object rotation angle of the target object at the reference moment relative to the first moment, so as to offset the change in object posture caused by the rotation.
[0114] Optionally, the image rotation angle may belong to the angle set {90, -90, 180, -180}. In this case, the server may determine the object rotation angle of the target object at the first moment relative to the reference moment based on the posture change of the target object in the reference image relative to the first image, and further select the image rotation angle whose opposite number is closest to the object rotation angle from the angle set. For example, Figure 9 As shown, the rotation angle of object A at the first moment relative to the reference moment is 80°. The corresponding image rotation angle at the first moment is -90°. The opposite of this image rotation angle, 90°, is closest to the object rotation angle of 80°. Since the photosensitive chip on the acquisition end and the screen on the playback end are typically rectangular, selecting an image rotation angle from the angle set eliminates the need to crop the rotated image, improving efficiency during the image update process.
[0115] In one possible implementation, Figure 9 As shown, the image metadata may include a frame index i for representing the position of the first image in the image sequence, gyroscope information gyro of the first terminal timestamp , core image information and image rotation angle r. The core image information may include feature position (x j ,y j), image width w and image height h.
[0116] In the above embodiment, determining the image metadata including the feature position and the image rotation angle can support various image adjustment methods such as image rotation and image translation, improve the flexibility of image processing, and further ensure image stability.
[0117] Optionally, in the process of processing images collected at multiple sampling moments, the reference moment can be a fixed sampling moment, such as the sampling moment specified by the user of the first terminal, or the first sampling moment; the reference moment can also be changed through loop iteration.
[0118] In an optional embodiment, the image processing method further includes: when the image metadata of the first moment is determined, using the first moment as a new reference moment, and using the updated image of the first moment as a new reference image; returning to the step of obtaining the first image captured by the first terminal at the first moment after the reference moment.
[0119] As mentioned above, if the sampling posture change at the first moment does not meet the change condition, or the reference image block and the first image block do not match, there is no need to adjust the image for the first moment. If the image metadata of the first moment is determined, it means that the first image captured at the first moment has undergone a change in the object posture relative to the reference image, and the first image needs to be adjusted according to the image metadata to obtain an updated image whose object posture matches the reference image. In this case, the first moment can be used as a new reference moment, and the updated image of the first moment can be used as a new reference image, and the process of obtaining the first image captured by the first terminal at the first moment after the reference moment can be returned to perform the next round of image processing.
[0120] It is understood that in the next round of image processing, the target object used to determine image metadata can be the same or different from the target object used in the previous round, thereby enabling adaptive adjustment of the image even when local content within the image changes. Furthermore, when the target object remains the same, dynamically changing the reference moment shortens the time interval between different sampling moments for object pose matching, thus avoiding significant changes in the object pose after adjustment and further ensuring image stability.
[0121] In actual applications, after determining the image metadata corresponding to the first image, adjustments can be made only to the first image, or at least one preceding image can be traced back and corresponding image metadata can be configured for the preceding image according to the image metadata corresponding to the first image. The adjustment amplitude of the preceding image is smaller than that of the first image, and, in the case of tracing back multiple preceding images, the closer the preceding image is to the first image in time sequence, the greater the adjustment amplitude. In this case, the first image and the preceding image of the first image can be adjusted based on the corresponding image metadata, so as to achieve a smooth transition of the image adjustment process, avoid jumps caused by image adjustment, and improve the visual effect during the subsequent image playback.
[0122] In an optional embodiment, the image metadata includes an image rotation angle. In this embodiment, the image processing method further includes: using the first moment as the rotation end moment, determining the rotation start moment from each sampling moment between the reference moment and the first moment; and determining the rotation angle corresponding to each selected sampling moment according to the time sequence of each selected sampling moment between the rotation start moment and the rotation end moment.
[0123] The time interval between the rotation start time and the rotation end time is positively correlated with the image rotation angle corresponding to the rotation end time. That is, the larger the image rotation angle corresponding to the first moment, the more previous images can be traced back.
[0124] Specifically, the server can use the first moment as the rotation end moment and determine the rotation start moment from each sampling moment between the reference moment and the first moment. Then, based on the temporal order of the selected sampling moments between the rotation start moment and the rotation end moment, the server determines the rotation angle corresponding to each selected sampling moment in a manner that increases with time.
[0125] Optionally, the server may determine the number of sampling moments between the reference moment and the first moment. Based on the image rotation angle corresponding to the first moment, the server may determine an expected number positively correlated with the image rotation angle. If the expected number is less than or equal to the number of sampling moments, the server may select a rotation start moment from each sampling moment according to the expected number. If the expected number is greater than the number of sampling moments, the server may determine the reference moment as the rotation start moment.
[0126] Optionally, the rotation angle difference between adjacent selected sampling moments can be the same; or the rotation angle corresponding to each selected sampling moment can be determined in a manner such that the rotation angle difference first decreases and then increases, so that the image adjustment process can present a slow-in and slow-out effect, further improving the user experience of the playback end.
[0127] It is understood that the rotation angle corresponding to any selected sampling moment is used to adjust the image captured at the selected sampling moment to obtain an updated image corresponding to the selected sampling moment. The image adjustment action can be performed by the server or the second terminal, and is not limited here.
[0128] In an optional embodiment, if Figure 10 As shown, the server can send the rotation start and end frames corresponding to the rotation start and end times as image metadata to the second terminal, so that the second terminal can rotate the acquired image information based on the rotation start and end frames. The rotation processing can be completed before or during the playback of the image. For example, when the frame index is 10, the rotation start and end frames are [5,10], and the image rotation angle is 50°, the second terminal can control the image to gradually rotate to 50° during the playback of the 5th to 10th frames.
[0129] When the sampling interval is short enough, the first terminal will collect multiple image frames during the process of sending the rotation, and the rotation angle of the same object in different image frames gradually increases. Based on this, by tracing back, the preceding image of the first image is also rotated, and the rotation angle gradually increases, which can achieve a smooth transition of the rotation in visual effects without any jumps, and can further enhance the user experience while ensuring image stability.
[0130] In one embodiment, Figure 11 As shown, an image processing method is provided, which can be executed by a computer device. Figure 1 Taking the second terminal in the example as an example, in this embodiment, the method includes the following steps:
[0131] Step S1102: Acquire image data corresponding to a plurality of consecutive sampling moments.
[0132] Optionally, the second terminal may obtain image data from the server. The image data may be sent by the first terminal to the server, or the server may further process the image information obtained from the first terminal to obtain image data and send it to the second terminal.
[0133] Step S1104 : for each sampling moment, when the image data corresponding to the sampling moment includes image metadata and image information, the image information is adjusted based on the image metadata to obtain an updated image corresponding to the sampling moment.
[0134] The image metadata is obtained based on the image processing method described in the above embodiment. Specifically, each image data may or may not include the corresponding image metadata. If the image data includes both image metadata and image information, the second terminal may adjust the image information based on the image metadata to obtain an updated image corresponding to the sampling moment.
[0135] Optionally, when the image metadata includes an image rotation angle, the second terminal may rotate the image represented by the image information according to the image rotation angle to obtain a corresponding updated image.
[0136] Optionally, if the image metadata includes a characteristic location, the second terminal may translate the image represented by the image information based on the characteristic position to obtain a corresponding updated image. Optionally, during image adjustment based on the characteristic location, the image block to which the characteristic location belongs may be moved, or the entire image may be translated.
[0137] Optionally, when the entire image is translated, edge images may be missing. Therefore, the server can enlarge and crop the translated image to ensure the integrity of the image.
[0138] Optionally, the process of adjusting the image information based on the image metadata may be performed before the image is played or during the image playback.
[0139] The above-mentioned image processing method obtains image data corresponding to multiple consecutive sampling moments; for each sampling moment, when the image data corresponding to the sampling moment includes image metadata and image information, the image information is adjusted based on the image metadata to obtain an updated image corresponding to the sampling moment. Because the image metadata is obtained through semantic segmentation and object pose recognition, the accuracy of the image metadata can be ensured. Therefore, regardless of how the posture parameters of the first terminal change, the image information collected at the sampling moment can be adjusted based on the image metadata corresponding to the sampling moment to obtain an updated image whose object pose matches the reference image. This allows the object pose of the same target object at different sampling moments to match, which is beneficial to improving image stability.
[0140] In an optional embodiment, the image metadata includes an image rotation angle. In this embodiment, for each sampling moment, when the image data corresponding to the sampling moment includes image metadata and image information, the image information is adjusted based on the image metadata to obtain an updated image corresponding to the sampling moment, including: determining, from each sampling moment, a target moment including the image metadata in the corresponding image data, and non-target moments other than the target moment; for each target moment, using the target moment as the rotation end moment, and determining a rotation start moment from the non-target moments between the target moment and the previous target moment; determining, based on the time sequence of each selected sampling moment between the rotation start moment and the rotation end moment, the rotation angle corresponding to each selected sampling moment; and adjusting the image information at each selected sampling moment based on each rotation angle to obtain an updated image corresponding to each selected sampling moment.
[0141] The time interval between the rotation start moment and the rotation end moment is positively correlated with the image rotation angle corresponding to the rotation end moment; the rotation angle increases sequentially over time. Specifically, the second terminal can parse each image data and determine, from each sampling moment, the target moment, including image metadata, and non-target moments other than the target moment, in the corresponding image data. Then, for each target moment, the second terminal can use the target moment as the rotation end moment and determine the rotation start moment from the non-target moments between the target moment and the previous target moment. Next, based on the temporal sequence of each selected sampling moment between the rotation start moment and the rotation end moment, the second terminal determines the rotation angle corresponding to each selected sampling moment, such that the rotation angle increases sequentially over time. Finally, the second terminal can adjust the image information for each selected sampling moment based on the rotation angles, obtaining an updated image corresponding to each selected sampling moment.
[0142] Optionally, the server may determine the number of sampling moments between the target moment and the previous target moment. Based on the image rotation angle corresponding to the target moment, the server may determine an expected number positively correlated with the image rotation angle. If the expected number is less than or equal to the number of sampling moments, the server may select a rotation start moment from each sampling moment according to the expected number. If the expected number is greater than the number of sampling moments, the server may determine the previous target moment as the rotation start moment.
[0143] Optionally, the rotation angle difference between adjacent selected sampling moments can be the same; or the rotation angle corresponding to each selected sampling moment can be determined in a manner such that the rotation angle difference first decreases and then increases, so that the image adjustment process can present a slow-in and slow-out effect, further improving the user experience of the playback end.
[0144] Optionally, after rotating the first image according to the image rotation angle, the rotated image may be further processed to adapt to the playback page size of the second terminal. The processing may include at least one of cropping or scaling.
[0145] like Figure 12 As shown, after the original image 1201 is rotated, the rotated image 1202 may exceed the play page 1203 of the second terminal. Figure 12 As shown, the second terminal can perform a zoom process on the rotated image 1202 to ensure that the image information in the original image 1201 can be fully displayed. The magnification factor used in the zoom process can be expressed as: scale = d4 / d3. Wherein, d4 is the width of the playback page of the second terminal, and d3 is the length of the original image 1201. Optionally, the second terminal can also crop the rotated image 1202, cutting off the image content 1204 and 1205 that exceeds the playback page to ensure that the played image matches the playback page. Optionally, as Figure 12 As shown, the missing part of the screen in the playback page after rotation can be filled with preset content or background color.
[0146] In a possible implementation, the second terminal may divide each continuous sampling moment into multiple sequences based on the target moment, and perform image adjustment for each sequence respectively, so as to synchronously process multiple image information, which is conducive to improving work efficiency.
[0147] In the above embodiment, by tracing back, the image collected at the preceding sampling moment of the target moment is also rotated, and the rotation angle gradually increases, which can achieve a smooth transition of rotation in visual effect without any jumps, and can further enhance the user experience while ensuring image stability.
[0148] In one embodiment, the image metadata includes image size and feature locations corresponding to the image information. In this embodiment, adjusting the image information based on the image metadata to obtain an updated image corresponding to the sampling time includes: translating the image represented by the image information so that the feature location is at a feature location on the playback page; and adjusting the image size based on the page size of the playback page to obtain an updated image with an image size that matches the playback page.
[0149] The characteristic position of the page can refer to the center of the page or the golden section point of the page. Optionally, the entire playback page can be divided into two regions in the width direction, where the ratio of the width of one region to the page width is equal to the ratio of the width of the other region to the width of the first region, and this ratio is approximately 0.618. Similarly, the entire playback page can be divided into two regions in the length direction, thereby obtaining two region dividing lines in the length and width directions. The intersection of these two region dividing lines can be used as the golden section point of the playback page.
[0150] Specifically, when the image metadata includes a characteristic position, the second terminal can, on the one hand, translate the image represented by the image information so that the characteristic position corresponding to the image information is at the page characteristic position of the playback page, ensuring that the image at the characteristic position can be highlighted. On the other hand, the second terminal can adjust the image size based on the page size of the playback page to obtain an updated image whose image size matches the playback page. Optionally, the way to adjust the image size may include zooming in, zooming out, cropping, etc. Figure 8 As shown, after the image translation is completed, the updated image matching the page size of the playback page can be obtained by zooming in and cropping the image. Wherein, d2 can refer to the distance between the page feature position and the page edge, which can be determined according to the page size and the page feature position.
[0151] In the above embodiment, when the image metadata includes feature positions and image sizes, the Hull image size can be adjusted by translating the image to obtain an updated image whose image size matches the playback page. This can ensure the compatibility of the updated image with the playback page, which is conducive to further improving the playback effect and prompting the user experience of the playback end.
[0152] In one embodiment, Figure 13 As shown, an image processing method is provided, which can be executed by a computer device, which can be a terminal or a server. Figure 1 Taking the server 130 in the example, in this embodiment, the method includes the following steps:
[0153] Step S1301: Acquire a reference image captured by a first terminal at a reference time, and an attitude angle corresponding to the first terminal at the reference time;
[0154] Step S1302: Acquire a first image captured by a first terminal at a first moment after a reference moment, and an attitude angle corresponding to the first terminal at the first moment;
[0155] Step S1303: determining a sampling posture change of the first terminal at the first moment relative to the reference moment based on the posture angles;
[0156] Step S1304: When the sampling posture change satisfies the change condition, semantic segmentation is performed on the reference image and the first image respectively to determine a plurality of reference candidate image blocks included in the reference image and a plurality of first candidate image blocks included in the first image;
[0157] Step S1305 , determining a plurality of reference image blocks belonging to the foreground image from each reference candidate image block, and determining a plurality of first image blocks belonging to the foreground image from each first candidate image block;
[0158] Step S1306 , determining the priority of each first image block according to the object category of the object depicted by each first image block;
[0159] Step S1307, determining a selected image block with the highest priority from among the first image blocks;
[0160] Step S1308, performing image feature matching on the selected image block and each reference image block respectively;
[0161] Step S1309: If there is an image block matching the selected image block among the reference image blocks, then the image block is determined as the target image block;
[0162] Step S1310: If all reference image blocks do not match the selected image block, then determine the selected image block with the highest priority from the remaining first image blocks; and return to step S1308;
[0163] Step S1311: If all reference image blocks do not match the selected image block and all first image blocks have been traversed, the first image is sent to the second terminal, and the next sampling moment of the first moment is used as the new first moment; and the process returns to step S1302;
[0164] Step S1312, determining the object position of the target object represented by the target image block in the first image as the feature position at the first moment;
[0165] Step S1313, determining an image rotation angle at the first moment relative to the reference moment based on the corresponding object postures of the target object in the reference image and the first image respectively;
[0166] Step S1314, determining image metadata including feature positions and image rotation angles;
[0167] The image metadata is used to adjust the first image to obtain an updated image in which the object pose matches the reference image;
[0168] Step S1315: taking the first moment as the rotation end moment, and determining the rotation start moment from each sampling moment between the reference moment and the first moment;
[0169] The time interval between the rotation start time and the rotation end time is positively correlated with the image rotation angle corresponding to the rotation end time.
[0170] Step S1316, determining the rotation angle corresponding to each selected sampling moment according to the time sequence of each selected sampling moment between the rotation start moment and the rotation end moment, and obtaining image metadata including the rotation angle;
[0171] Among them, the rotation angle increases sequentially with time;
[0172] Step S1317: Use the first moment as a new reference moment, and use the updated image at the first moment as a new reference image; and return to step S1302;
[0173] Step S1318: sending the image metadata and image information to the second terminal according to the sampling time;
[0174] The image metadata corresponding to any sampling moment is used to adjust the image collected at the sampling moment to obtain an updated image corresponding to the sampling moment;
[0175] Step S1319: When the sampling posture change does not satisfy the change condition, the first image is sent to the second terminal, and the next sampling moment of the first moment is used as the new first moment; and the process returns to step S1302.
[0176] Optionally, the second terminal can continuously obtain image data corresponding to continuous sampling moments from the server. For each sampling moment, if the image data at the sampling moment contains image information but does not contain corresponding image metadata, the second terminal can play the image represented by the image information. If the image data at the sampling moment contains image metadata and image information, the second terminal can adjust the image represented by the image information based on the image metadata, obtain an updated image and play it. Therefore, during the continuous playback of the image, the stability of the object posture of the same object in the image can always be maintained, thereby ensuring the stability of the image.
[0177] The above-mentioned image processing method obtains image metadata through semantic segmentation and object posture recognition, which can ensure the accuracy of the image metadata. Therefore, no matter how the posture parameters of the first terminal change, the image information collected at the sampling moment can be adjusted according to the image metadata corresponding to the sampling moment, and an updated image in which the object posture matches the reference image is obtained, so that the object posture of the same target object at different sampling moments can be matched, which is conducive to improving the stability of the image.
[0178] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0179] Based on the same inventive concept, embodiments of the present application also provide an image processing device for implementing the aforementioned image processing method. The solution provided by this device is similar to the solution described in the aforementioned method. Therefore, the specific limitations in one or more of the following image processing device embodiments can be found in the above-described limitations on the image processing method and will not be further elaborated here.
[0180] In one embodiment, Figure 14 As shown, an image processing device is provided, including: an image acquisition module 1401, a semantic segmentation module 1402 and a metadata determination module 1403, wherein:
[0181] The image acquisition module 1401 is configured to acquire a reference image captured by the first terminal at a reference time and a first image captured by the first terminal at a first time after the reference time;
[0182] A semantic segmentation module 1402 is configured to perform semantic segmentation on the reference image and the first image based on the posture parameters corresponding to the first terminal at the reference time and the first time, respectively, to determine a reference image block included in the reference image and a first image block included in the first image;
[0183] The metadata determination module 1403 is used to obtain a target image block that matches the first image block from the reference image block, and determine the image metadata at the first moment based on the object pose corresponding to the target object represented by the target image block in the reference image and the first image respectively; the image metadata is used to adjust the first image to obtain an updated image whose object pose matches the reference image.
[0184] In one embodiment, the semantic segmentation module 1402 includes: a posture parameter acquisition unit, used to obtain the posture parameters corresponding to the first terminal at the reference moment and the first moment respectively; a posture change determination unit, used to determine the sampling posture change of the first terminal at the first moment relative to the reference moment based on each posture parameter; a semantic segmentation unit, used to perform semantic segmentation on the reference image and the first image respectively when the sampling posture change meets the change condition.
[0185] In one embodiment, the attitude parameter includes an attitude angle. In this embodiment, the attitude parameter acquisition unit is specifically configured to: acquire gyroscope information corresponding to the first terminal at the reference time and the first time; and determine the attitude angle of the first terminal based on the gyroscope information.
[0186] In one embodiment, the semantic segmentation module 1402 is specifically used to: perform semantic segmentation on the reference image and the first image respectively, determine multiple reference candidate image blocks contained in the reference image, and multiple first candidate image blocks contained in the first image; determine the reference image block belonging to the foreground image from each reference candidate image block; and determine the first image block belonging to the foreground image from each first candidate image block.
[0187] In one embodiment, there are multiple reference image blocks and multiple first image blocks. In this embodiment, the image processing device further includes an image feature matching module, which is configured to: determine the priority of each first image block based on the object category of the object depicted by each first image block; determine a selected image block with the highest priority from each first image block, and perform image feature matching on the selected image block with each reference image block; if there is an image block in each reference image block that matches the selected image block, determine the image block as the target image block; if none of the reference image blocks matches the selected image block, determine the selected image block with the highest priority from the remaining first image blocks, and return to the step of performing image feature matching on the selected image block with each reference image block.
[0188] In one embodiment, the metadata determination module is specifically used to: determine the object position of the target object represented by the target image block in the first image as the feature position at the first moment; determine the image rotation angle at the first moment relative to the reference moment based on the object postures corresponding to the target object in the reference image and the first image respectively; and determine image metadata including the feature position and the image rotation angle.
[0189] In one embodiment, the image processing device also includes an iteration module for: when the image metadata of the first moment is determined, taking the first moment as a new reference moment, and taking the updated image of the first moment as a new reference image, and returning to the step of obtaining the first image captured by the first terminal at the first moment after the reference moment.
[0190] In one embodiment, the image metadata includes an image rotation angle; in the case of this embodiment, the metadata determination module 1403 is further used to: take the first moment as the rotation end moment, and determine the rotation start moment from each sampling moment between the reference moment and the first moment; determine the rotation angle corresponding to each selected sampling moment according to the time sequence of each selected sampling moment between the rotation start moment and the rotation end moment; wherein the time interval between the rotation start moment and the rotation end moment is positively correlated with the image rotation angle corresponding to the rotation end moment; the rotation angle increases sequentially with time; the rotation angle corresponding to any selected sampling moment is used to adjust the image captured at the selected sampling moment to obtain an updated image corresponding to the selected sampling moment.
[0191] In one embodiment, Figure 15 As shown, an image processing device is provided, including: an image data acquisition module 1501 and an image adjustment module 1502, wherein:
[0192] The image data acquisition module 1501 is used to acquire image data corresponding to a plurality of consecutive sampling moments;
[0193] The image adjustment module 1502 is used to adjust the image information based on the image metadata for each sampling moment, when the image data corresponding to the sampling moment includes image metadata and image information, to obtain an updated image corresponding to the sampling moment; the image metadata is obtained based on the above-mentioned image processing method.
[0194] In one embodiment, the image metadata includes an image rotation angle. In this embodiment, the image adjustment module 1502 is specifically configured to: determine, from each sampling moment, a target moment including the image metadata in the corresponding image data, as well as non-target moments other than the target moment; for each target moment, use the target moment as the rotation end moment, and determine a rotation start moment from the non-target moments between the target moment and the previous target moment; determine the rotation angle corresponding to each selected sampling moment according to the temporal sequence of each selected sampling moment between the rotation start moment and the rotation end moment; and adjust the image information of each selected sampling moment based on the rotation angle to obtain an updated image corresponding to each selected sampling moment. The time interval between the rotation start moment and the rotation end moment is positively correlated with the image rotation angle corresponding to the rotation end moment; and the rotation angle increases sequentially over time.
[0195] In one embodiment, the image metadata includes the image size and characteristic position corresponding to the image information. In this embodiment, the image adjustment module 1502 is specifically configured to: translate the image represented by the image information so that the characteristic position is at the characteristic position of the playback page; and adjust the image size based on the page size of the playback page to obtain an updated image whose image size matches the playback page.
[0196] Each module in the above-mentioned image processing device can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in the computer device in the form of software, so that the processor can call and execute the corresponding operations of each module.
[0197] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 16 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O) and a communication interface. The processor, memory and input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store data involved in the above-mentioned image processing process. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, an image processing method is implemented.
[0198] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 17As shown. The computer device includes a processor, memory, an input / output interface, a communication interface, a display unit, and an input device. The processor, memory, and input / output interface are connected via a system bus, and the communication interface, display unit, and input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal via wired or wireless communication, and the wireless communication can be achieved via WIFI, a mobile cellular network, NFC (near field communication), or other technologies. When the computer program is executed by the processor, it implements an image processing method. The display unit of the computer device is used to form a visually visible image, and can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse, etc.
[0199] Those skilled in the art will understand that Figure 16 The structure shown in or 17 is merely a block diagram of a portion of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0200] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps of the above method embodiment when executing the computer program.
[0201] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the above method embodiment are implemented.
[0202] In one embodiment, a computer program product is provided, comprising a computer program, which implements the steps of the above method embodiment when executed by a processor.
[0203] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of such data must comply with the relevant laws, regulations, and standards of the relevant regions and areas. Furthermore, the user may choose not to authorize the use of such information and related data, or may refuse or conveniently refuse to receive push notifications.
[0204] In this application, when collecting and processing relevant data in actual applications, the requirements of relevant local laws and regulations should be strictly followed to obtain the informed consent or separate consent of the personal information subject, and subsequent data use and processing should be carried out within the scope of authorization of laws and regulations and the personal information subject.
[0205] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. In particular, any reference to memory, database, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The databases involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided herein may be, but are not limited to, general-purpose processors, central processing units (CPUs), graphics processors (GPUs), digital signal processors (DSPs), programmable logic devices (PLCs), and the like.
[0206] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0207] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a reference image captured by a first terminal at a reference time, and a first image captured by the first terminal at a first time after the reference time; performing semantic segmentation on the reference image and the first image based on posture parameters corresponding to the first terminal at the reference moment and the first moment, respectively, to determine a reference image block included in the reference image and a first image block included in the first image; A target image block matching the first image block is obtained from the reference image block, and image metadata at the first moment is determined based on the object poses of the target object represented by the target image block in the reference image and the first image, respectively. The image metadata is used to adjust the first image to obtain an updated image whose object pose matches the reference image.
2. The method according to claim 1, characterized in that The performing semantic segmentation on the reference image and the first image based on the posture parameters corresponding to the first terminal at the reference time and the first time, respectively, includes: Obtaining posture parameters of the first terminal corresponding to the reference time and the first time respectively; determining, based on the posture parameters, a sampling posture change of the first terminal at the first moment relative to the reference moment; When the sampling posture change satisfies a change condition, semantic segmentation is performed on the reference image and the first image respectively.
3. The method according to claim 2, characterized in that The attitude parameter includes an attitude angle; and obtaining the attitude parameters of the first terminal corresponding to the reference time and the first time respectively includes: Obtaining gyroscope information corresponding to the first terminal at the reference time and the first time respectively; Determine an attitude angle of the first terminal based on the gyroscope information.
4. The method according to claim 1, wherein The performing semantic segmentation on the reference image and the first image respectively to determine the reference image block included in the reference image and the first image block included in the first image includes: Performing semantic segmentation on the reference image and the first image respectively to determine a plurality of reference candidate image blocks included in the reference image and a plurality of first candidate image blocks included in the first image; Determining a reference image block belonging to a foreground image from each of the reference candidate image blocks; A first image block belonging to a foreground image is determined from each of the first candidate image blocks.
5. The method according to claim 1, wherein The number of the reference image blocks and the number of the first image blocks are both multiple; the method further includes: determining a priority of each of the first image blocks according to an object category of an object depicted by each of the first image blocks; determining a selected image block with the highest priority from each of the first image blocks, and performing image feature matching between the selected image block and each of the reference image blocks; If there is an image block matching the selected image block in each of the reference image blocks, determining the image block as the target image block; If none of the reference image blocks matches the selected image block, a selected image block with the highest priority is determined from the remaining first image blocks, and the process returns to the step of performing image feature matching between the selected image block and the reference image blocks.
6. The method according to claim 1, wherein The determining of the image metadata at the first moment based on the object poses corresponding to the target object represented by the target image block in the reference image and the first image, respectively, includes: determining an object position of a target object represented by the target image block in the first image as a feature position at the first moment; determining an image rotation angle at the first moment relative to the reference moment based on the object postures corresponding to the target object in the reference image and the first image respectively; Image metadata including the feature location and the image rotation angle is determined.
7. The method according to any one of claims 1 to 6, characterized in that The method further comprises: When the image metadata at the first moment is determined, taking the first moment as a new reference moment, and taking the updated image at the first moment as a new reference image; Return to the step of acquiring a first image captured by the first terminal at a first moment after the reference moment.
8. The method according to claim 7, characterized in that The image metadata includes an image rotation angle; and the method further includes: The first moment is used as the rotation end moment, and a rotation start moment is determined from each sampling moment between the reference moment and the first moment; the time interval between the rotation start moment and the rotation end moment is positively correlated with the image rotation angle corresponding to the rotation end moment; According to the time sequence of each selected sampling moment between the rotation start moment and the rotation end moment, the rotation angle corresponding to each selected sampling moment is determined; the rotation angle increases sequentially with the time sequence; the rotation angle corresponding to any of the selected sampling moments is used to adjust the image collected at the selected sampling moment to obtain an updated image corresponding to the selected sampling moment.
9. An image processing method, characterized in that: The method comprises: Obtaining image data corresponding to a plurality of consecutive sampling moments; For each of the sampling moments, when the image data corresponding to the sampling moment includes image metadata and image information, the image information is adjusted based on the image metadata to obtain an updated image corresponding to the sampling moment; the image metadata is obtained based on the method described in any one of claims 1 to 8.
10. The method according to claim 9, characterized in that The image metadata includes an image rotation angle; and for each sampling moment, when the image data corresponding to the sampling moment includes the image metadata and image information, adjusting the image information based on the image metadata to obtain an updated image corresponding to the sampling moment, comprising: Determining, from each of the sampling moments, a target moment including image metadata in the corresponding image data, and non-target moments other than the target moment; For each target moment, the target moment is used as the rotation end moment, and a rotation start moment is determined from non-target moments between the target moment and the previous target moment; the time interval between the rotation start moment and the rotation end moment is positively correlated with the image rotation angle corresponding to the rotation end moment; Determining the rotation angles corresponding to the selected sampling moments according to a time sequence of the selected sampling moments between the rotation start moment and the rotation end moment; wherein the rotation angles increase sequentially with the time sequence; Based on the rotation angles, the image information of each selected sampling moment is adjusted to obtain an updated image corresponding to each selected sampling moment.
11. The method according to claim 9, characterized in that The image metadata includes an image size and a feature position corresponding to the image information; and adjusting the image information based on the image metadata to obtain an updated image corresponding to the sampling moment includes: translating the image represented by the image information so that the characteristic position is located at a page characteristic position of the playback page; The image size is adjusted based on the page size of the playing page to obtain an updated image whose image size matches the playing page.
12. An image processing device, characterized in that: The device comprises: an image acquisition module, configured to acquire a reference image acquired by a first terminal at a reference time, and a first image acquired by the first terminal at a first time after the reference time; a semantic segmentation module, configured to perform semantic segmentation on the reference image and the first image based on the posture parameters corresponding to the first terminal at the reference moment and the first moment, respectively, to determine a reference image block included in the reference image and a first image block included in the first image; a metadata determination module configured to obtain, from the reference image block, a target image block that matches the first image block, and determine image metadata at the first moment based on the object poses of a target object represented by the target image block in the reference image and the first image, respectively; the image metadata being used to adjust the first image to obtain an updated image whose object pose matches the reference image.
13. An image processing device, characterized in that: The device comprises: An image data acquisition module is used to acquire image data corresponding to a plurality of consecutive sampling moments; An image adjustment module is configured to adjust the image information based on the image metadata for each of the sampling moments, when the image data corresponding to the sampling moment includes image metadata and image information, so as to obtain an updated image corresponding to the sampling moment; the image metadata is obtained based on the method described in any one of claims 1 to 8.
14. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 11 are implemented.
15. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.
16. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.