Image semantic adaptive compression transmission method and system for mobile device
By analyzing inertial motion and interactive operation data, quantifying environmental interference and user intent, and adopting a differentiated coding strategy, the problem of balancing global smoothness and local operability in image transmission on mobile devices was solved, achieving a simultaneous improvement in user experience and transmission efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NORTHWEST UNIV
- Filing Date
- 2026-03-09
- Publication Date
- 2026-06-02
AI Technical Summary
Existing image transmission technologies cannot simultaneously guarantee user experience and transmission efficiency in mobile devices, especially under conditions of device shaking and complex user interaction, and cannot achieve a balance between global smoothness and local operability.
By simultaneously analyzing inertial motion data and image interaction data, the degree of environmental interference and user intent are quantified. A differentiated coding strategy is adopted to predict and delineate the interaction focus area. Resource allocation is optimized through historical interaction data to form an adaptive closed-loop logic.
In complex motion environments, it achieves the synchronization of global smoothness and local operability in image transmission, improves the balance between user experience and transmission efficiency, and avoids the problems of sudden drop in image quality or stuttering in traditional methods.
Smart Images

Figure CN121814908B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image transmission and communication technology, and more specifically, this application relates to an image semantic adaptive compression transmission method and system for mobile devices. Background Technology
[0002] Image transmission on mobile devices in complex motion environments faces an inherent contradiction between user experience and transmission efficiency. When users use their devices while moving, the physical jitter of the device and the frequent user interactions intertwine, posing a dual challenge to the image quality and smoothness of real-time image services. Existing image compression and transmission technologies typically employ a globally uniform compression strategy, which struggles to cope with such dynamic and heterogeneous perception requirements. This leads to situations in bandwidth-constrained mobile networks where either excessive compression compromises the readability of critical information, or stuttering and high latency occur to maintain image quality, making it difficult to achieve a satisfactory user experience.
[0003] Specifically, existing solutions largely rely on passive adaptation to a single dimension: network bandwidth or device motion. For example, when severe device shaking is detected, a common approach is to increase the compression ratio across the board to maintain smooth transmission. However, this method ignores the user's real-time interactive intent: even during overall shaking, the user's visual attention and operational goals are usually focused on specific areas of the screen, such as a point of interest on a map or a section of text in a document. Over-compression of these critical areas directly prevents users from performing accurate reading, recognition, or clicking operations, resulting in inefficient allocation of bandwidth resources.
[0004] Furthermore, some technologies that attempt to introduce region-of-attention recognition often rely on static or semi-static analysis of image content, lacking real-time, continuous response to user interaction intentions. A user's focus of attention dynamically shifts with their operational trajectory, such as from rapid swiping to slowing down and hovering, and then to precise selection—a continuously changing process. Traditional methods cannot predict a user's intended focus before they begin precise operations and cannot reserve high-quality transmission resources in advance. Consequently, image quality often drops sharply or response delays occur at key interaction frames, disrupting the continuity and anticipation of the interaction process. This lag in perception and decision-making means the system is always one step behind the user's actual needs.
[0005] To address the aforementioned issues, there is an urgent need in this field for a method that can understand and predict user interaction intentions in real time under the background of device motion interference, thereby enabling proactive and refined control of image transmission resources, in order to resolve the core contradiction that global smoothness and local operability of image transmission cannot be guaranteed simultaneously in mobile scenarios. Summary of the Invention
[0006] To address the aforementioned technical problems, this paper provides a method and system for image semantic adaptive compression transmission for mobile devices. This technical solution solves the problems mentioned in the background section.
[0007] In a first aspect, embodiments of this application provide an image semantic adaptive compression transmission method for mobile devices, comprising the following steps: acquiring the original inertial motion data and original image interaction operation data of the target device in the current period, performing time-frequency domain transformation and trajectory extrapolation analysis respectively to obtain an interference image index and an interaction operation fitting piecewise linear curve; if the interference image index is greater than a preset motion interference threshold, mapping the interference image index to a global base compression ratio, and obtaining several target region points by dividing the interaction operation fitting piecewise linear curve according to a preset segmentation rule; delineating the interaction focus region of the target region points with a preset radius to obtain several interaction focus regions, and calculating the calibration based on the ratio of the intersection area of the several interaction focus regions to the sum of the areas of the interaction focus regions. The coefficients are used to reduce the global base compression ratio to obtain the first compression parameter, and the coefficients are used to increase the global base compression ratio to obtain the second compression parameter. Based on the first compression parameter and the second compression parameter, differential encoding is performed on the interactive focus area and the area outside the area to obtain the first area encoding and the second area encoding, and these are merged into a complete frame bitstream for transmission. The fitted polyline of the interactive operation within a preset time window containing at least two historical periods is obtained, the target area point sequence is extracted and spatial two-dimensional kernel density is estimated accordingly, the local density value of the target area points is obtained, and the preset radius is corrected accordingly to obtain the corrected radius. The interactive focus area is then redefined and used for the next period.
[0008] Secondly, embodiments of this application provide an image semantic adaptive compression transmission system for mobile devices, comprising: a data acquisition module for acquiring the original inertial motion data and original image interaction operation data of the target device in the current period, performing time-frequency domain transformation and trajectory extrapolation analysis respectively to obtain an interference image index and an interaction operation fitting piecewise linear curve; a target region point acquisition module for mapping the interference image index to a global base compression ratio if the interference image index is greater than a preset motion interference threshold, and obtaining several target region points from the interaction operation fitting piecewise linear curve according to a preset segmentation rule; and a calibration coefficient acquisition module for delineating the interaction focus region of the target region points with a preset radius to obtain several interaction focus regions, and calculating the calibration coefficient based on the ratio of the intersection area of the several interaction focus regions to the sum of the areas of the interaction focus regions; compression... The compression parameter acquisition module is used to reduce the global base compression ratio according to the adjustment coefficient to obtain the first compression parameter, and increase the global base compression ratio according to the adjustment coefficient to obtain the second compression parameter. The image encoding and transmission module is used to perform differential encoding on the interactive focus area and the area outside it according to the first compression parameter and the second compression parameter to obtain the first region encoding and the second region encoding, and merge them into a complete frame bitstream for transmission. The interactive focus area re-delineation module is used to obtain the interactive operation fitting polyline within a preset time window containing at least two historical periods, extract the target area point sequence and perform spatial two-dimensional kernel density estimation based on it to obtain the local density value of the target area points, and correct the preset radius accordingly to obtain the corrected radius, and redefine the interactive focus area for the next period.
[0009] Thirdly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described image semantic adaptive compression transmission method for mobile devices.
[0010] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0011] 1. By simultaneously analyzing raw inertial motion data and raw image interaction data, and converting them into interference image indices and interaction operation fitting line graphs respectively, this approach can simultaneously quantify the objective impact of the external motion environment on image readability, as well as the user's subjective attention direction and operation trends. This dual-dimensional perception mechanism ensures that subsequent resource allocation decisions are no longer based solely on network bandwidth or device jitter, but rather on a dynamic coupling analysis of "environmental interference level" and "user intent focus." This fundamentally solves the problem that existing technologies, due to their single perception dimension, cannot intelligently guarantee image quality in key interactive areas under shaky conditions.
[0012] 2. Instead of uniformly adjusting the quality of the entire frame, the interactive focus area is defined based on the target area points predicted by the fitted polyline of the interactive operation. Then, fine-tuning is performed on the global base compression ratio in the opposite direction using adjustment coefficients, generating differentiated first and second compression parameters. Essentially, this dynamically redistributes valuable encoded bit resources from background areas unrelated to the user's attention to their focus area. This "spatial resource scheduling" strategy, guided by interactive semantics, can reserve high-quality encoded resources for the user's attention area before they perform precise clicks or reading, effectively solving the problem of "sudden drop in image quality or response delay in interactive keyframes" mentioned in the background technology, ensuring a smooth and clear interactive process.
[0013] 3. By introducing a spatial two-dimensional kernel density estimation based on fitted polygonal lines of historical cycle interaction operations, the spatial clustering characteristics of user operation habits can be analyzed, and the range of the interaction focus area for the next cycle can be corrected accordingly. Furthermore, the prediction and correction of inertial slip offset data in inertial compensation processing further refines the judgment of the user's intended landing point. These mechanisms not only respond to the current state but also learn from continuous interactions and optimize future decisions, forming a continuously evolving adaptive closed loop. This solves the problem that static or reactive ROI encoding cannot adapt to the continuous migration of user attention, making resource allocation strategies more accurate and intelligent, keeping pace with the user's dynamically changing intentions, and ultimately achieving a lasting optimal balance between user experience and transmission efficiency in complex mobile usage scenarios. Attached Figure Description
[0014] Figure 1 This is a schematic diagram illustrating the steps of an image semantic adaptive compression transmission method for mobile devices provided in an embodiment of this application.
[0015] Figure 2 A schematic diagram of the logical flow of the image semantic adaptive compression transmission method for mobile devices provided in the embodiments of this application;
[0016] Figure 3 This is a schematic diagram of the structure of an image semantic adaptive compression transmission system for mobile devices provided in an embodiment of this application. Detailed Implementation
[0017] This application's embodiments address the technical problem in the prior art where there is insufficient real-time understanding and prediction of user interaction intentions in image transmission in mobile scenarios, thus hindering the fine-grained control of image transmission resources, through an image semantic adaptive compression transmission method and system for mobile devices.
[0018] This approach stems from a clear real-world contradiction: in mobile device use, there is a direct conflict between image blurring caused by motion and maintaining the clarity of interactive operations. Traditional solutions treat these two as an irreconcilable contradiction, resorting to a globally uniform compression strategy to make trade-offs, often resulting in the sacrifice of one. The fundamental idea behind this approach is to recognize that the user's real need is not absolute clarity across the entire screen, but rather ensuring the clarity and usability of specific local areas intended for interaction under the objective condition of device movement. Therefore, the core problem becomes how to accurately and in real-time locate and protect these local areas in a dynamic environment, and use the resulting saved encoding resources to compensate for the additional bandwidth overhead incurred in maintaining clarity.
[0019] Based on this, this solution establishes the underlying logic of two parallel data streams: collaborative perception and processing. First, it continuously analyzes the device's raw inertial motion data, quantifying it into an interference image index that characterizes the risk of impaired overall image readability. This defines the objective environmental conditions for initiating a global bandwidth-saving strategy. Second, it analyzes the user's raw image interaction data in real time, forming an interaction operation fitting polygonal line by analyzing its continuous trajectory, thereby extracting and predicting the target area point representing the user's attention. When motion interference is significant, this solution couples these two information streams for decision-making: on the one hand, it sets a global base compression ratio applicable to the entire image based on the degree of interference, serving as a baseline for bandwidth control; on the other hand, it delineates the corresponding interaction focus area in the image based on the target point predicted by the interaction intent. Crucially, this solution does not simply apply a fixed high-quality encoding to the focus area independent of the environment, but introduces a calibration coefficient reflecting the spatial distribution relationship of multiple focus areas. This coefficient determines whether the user's attention is scattered or concentrated by analyzing the intersection of each area. If the areas are scattered, the improvement in image quality for each area is moderately suppressed to prevent excessive resource dispersion; if the areas are highly overlapping, the image quality guarantee for the core overlapping areas is enhanced. Ultimately, this scheme uses tuning coefficients to differentiate the global base compression ratio, generating a first compression parameter specifically for the focal region and a second compression parameter for the remaining background regions. The former has higher quality, while the latter is more aggressive in compression, thereby redistributing a limited number of coding bits in the spatial dimension and achieving local operability assurance under global fluency constraints.
[0020] To enhance the forward-looking and adaptive nature of this solution, an optimization mechanism based on historical behavior learning is further introduced. This solution collects historical interaction data and analyzes the spatial distribution density of point sequences in the target area to obtain local density values reflecting the degree of clustering of user operating habits. Based on this, the predicted area size is dynamically adjusted to form a more accurate correction radius for subsequent cycles. Furthermore, for the common operation of a sudden stop after a rapid swipe, this solution analyzes speed changes and trajectory curvature to determine if the user made any final directional adjustments and dynamically corrects the landing point offset based on inertial prediction, ensuring that the defined focal area more accurately matches the user's true intended landing point.
[0021] In summary, the core innovation of this solution lies in establishing a closed-loop data processing logic of "environmental interference quantification - interaction intent prediction - resource space reallocation - continuous behavior learning." It doesn't invent a new encoding algorithm, but rather creatively integrates the two independent dimensions of device motion sensing and user interaction analysis, driving the encoder to perform intelligent and proactive bitrate space allocation. This enables the solution to intelligently distinguish the "target area" where the user is about to perform a precise operation from the "background area" where smoothness is only needed, even in highly interactive scenarios such as navigation and mobile reading, where device shaking is inevitable. This breakthrough solves the traditional dilemma of "blurring key operation points to ensure overall smoothness" or "causing global lag to ensure operation point clarity," achieving for the first time real-time, accurate response and resource guarantee to user interaction intent within limited mobile network bandwidth.
[0022] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0023] Figure 1This is a schematic diagram illustrating the steps of the image semantic adaptive compression transmission method for mobile devices provided in this application embodiment. The image semantic adaptive compression transmission method for mobile devices includes the following steps: acquiring the original inertial motion data and original image interaction operation data of the target device in the current cycle, performing time-frequency domain transformation and trajectory extrapolation analysis respectively to obtain the interference image index and the interaction operation fitting piecewise linear curve; if the interference image index is greater than a preset motion interference threshold, mapping the interference image index to a global base compression ratio, and obtaining several target region points by dividing the interaction operation fitting piecewise linear curve according to a preset segmentation rule; delineating the interaction focus region of the target region points with a preset radius to obtain several interaction focus regions, and determining the interaction focus region based on the intersection area of the several interaction focus regions and the interaction focus region surface. The ratio of the sum of products is used to calculate the adjustment coefficient; the global base compression ratio is reduced according to the adjustment coefficient to obtain the first compression parameter, and the global base compression ratio is increased according to the adjustment coefficient to obtain the second compression parameter; differential encoding is performed on the interactive focus area and the area outside the area according to the first compression parameter and the second compression parameter to obtain the first area code and the second area code, and these are merged into a complete frame bitstream for transmission; the interactive operation fitting polyline within a preset time window containing at least two historical periods is obtained, the target area point sequence is extracted and spatial two-dimensional kernel density is estimated accordingly to obtain the local density value of the target area points, and the preset radius is corrected accordingly to obtain the corrected radius, and the interactive focus area is redefined accordingly for the next period.
[0024] The preset radius is obtained by: calculating the difference between the interference image index and the preset motion interference threshold; multiplying the difference by a preset radius scaling factor to obtain a radius adjustment amount; and adding the radius adjustment amount to a preset base radius value. The result is the preset radius used to delineate the interactive focus area.
[0025] Figure 2This is a schematic diagram of the logical flow of the image semantic adaptive compression transmission method for mobile devices provided in the embodiments of this application. The present invention creatively solves the core contradiction of the difficulty in balancing the global smoothness and local operability of image transmission on mobile devices under motion interference through the closed-loop data processing logic of "environmental interference quantification - interaction intent prediction - resource space reallocation - behavior continuous learning". Its technical effects are as follows: First, this embodiment, by synchronously analyzing device inertial motion data and user interaction operation data, deeply integrates the objective "interference image index" with the "interaction operation fitting piecewise linear curve" representing subjective intent, realizing a shift from single-environment adaptation to "environment-intent" dual-dimensional collaborative decision-making; Second, it innovatively introduces a "calibration coefficient" reflecting the spatial overlap relationship of multiple interaction focus areas, and accordingly makes differentiated adjustments to the global basic compression ratio, generating a low compression ratio "first compression parameter" for focus areas and a high compression ratio "second compression parameter" for background areas, thereby realizing intelligent scheduling of encoded bits from non-interested areas to interest areas at the pixel space level, significantly improving the visual clarity of the intent interaction area while ensuring overall transmission smoothness; Finally, this embodiment obtains "local density values" by analyzing the spatial density of historical interaction points to dynamically correct the predicted area range, and combines inertial offset compensation for sliding emergency stop operations to build an adaptive capability of continuous learning and forward optimization, enabling resource allocation to continuously and accurately match the user's dynamically changing intent. Ultimately, this solution achieves, for the first time under limited bandwidth, the real-time and intelligent flow of network resources in accordance with the user's visual attention and operational intentions, completely breaking through the inherent dilemma of traditional global compression strategies in shaking scenarios that "either blur key points or cause global lag".
[0026] Furthermore, the specific process for obtaining the adjustment coefficient is as follows: the ratio of the intersection area of several interactive focus areas to the sum of the areas of the interactive focus areas is recorded as the first ratio; the intersection area of interactive focus areas whose first ratio is less than the first preset ratio is counted as the first quantity; the intersection area of interactive focus areas whose first ratio is greater than or equal to the first preset ratio is counted as the second quantity; when the first quantity is greater than or equal to the second quantity, it is determined that the interactive focus areas are dispersed, and the first preset coefficient is multiplied by the first preset magnification factor to obtain the adjustment coefficient; when the second quantity is greater than the first quantity, it is determined that the interactive focus areas are overlapping, and the first preset coefficient is multiplied by the first preset reduction factor to obtain the adjustment coefficient.
[0027] In this embodiment, the first preset ratio is used to determine the degree of overlap between a single interactive focus area and other areas, thereby distinguishing whether the area is relatively independent or highly overlapping. The setting is based on the following: when the percentage of the intersection area between areas is less than this value, the user's attention in that area is considered relatively independent; if it exceeds or equals this value, it is considered likely to be highly related to other points of interest. A typical value range is 0.2 to 0.4, for example, it can be set to 0.3.
[0028] First preset coefficient: This serves as the base for calculating the calibration coefficient. It can usually be set to the number 1, representing the baseline state without calibration.
[0029] First preset magnification factor: When the interactive focus area is determined to be scattered, this is the multiplier used to amplify the adjustment coefficient. The rationale for this setting is that when areas are scattered, the user's attention is relatively evenly distributed. Therefore, this embodiment should provide a relatively conservative improvement in image quality for each area to avoid excessive resource dispersion leading to insignificant improvements in each area. Thus, this factor should be greater than 1, for example, between 1.2 and 1.5.
[0030] First preset reduction factor: When the interactive focus areas are determined to be overlapping, this is used as a multiplier to reduce the adjustment coefficient. The rationale for this setting is that highly overlapping areas indicate a high level of user focus, and this embodiment should enhance the image quality improvement of the core overlapping areas. Therefore, this factor should be less than 1, for example, between 0.6 and 0.8.
[0031] Count the first and second quantities: This operation is based on the counting logic, which traverses all interactive focus areas and accumulates the count of the two types of areas according to the comparison result of their "first ratio" and "first preset ratio".
[0032] The spatial topological relationships between interactive focus areas are transformed into a quantifiable control signal. It doesn't directly process image content, but rather performs meta-analysis on the set of interactive focus areas representing user intent. When multiple areas overlap, it means that multiple user actions or predictions point to similar locations on the screen, strongly suggesting that this is the absolute core of the current task. Therefore, a first preset reduction factor is used to increase the adjustment coefficient, resulting in a smaller and higher-quality first compression parameter, and a larger second compression parameter for more aggressive background compression, concentrating resources more on the core area. Conversely, scattered areas suggest that the user may be browsing or comparing; in this case, uniformly and moderately improving the image quality of each area is more reasonable.
[0033] Through the above technical solution, this implementation dynamically determines whether the user's attention is focused or scattered by analyzing the distribution of the intersection area ratio of each interactive focus area, and adaptively adjusts the global compression intensity accordingly. This achieves intelligent allocation of encoding resources based on the spatial aggregation degree of the user's attention, providing significant core area image quality improvement when attention is focused, and avoiding ineffective bandwidth dissipation when attention is scattered.
[0034] Furthermore, taking any target region point in the target region point sequence as the center, a circular neighborhood is defined with a preset radius. The weighted contribution of other target region points within this circular neighborhood to the target region point is processed by spatial two-dimensional kernel density estimation to obtain the local density value of the target region point. The local density values of all target region points are normalized to obtain the density parameter of the target region point. If the density parameter of the target region point is less than or equal to the preset density threshold, the preset radius is not corrected. If the density parameter of the target region point is greater than the preset density threshold, a correction factor is obtained by adding the preset density influence coefficient and the density parameter to the number 1. The preset radius is multiplied by the correction factor to obtain the corrected radius of the target region point.
[0035] In this embodiment, a preset density influence coefficient is used: this coefficient controls the intensity of the correction of the radius of the current cycle's interaction focus area by the local density of historical operation points. The coefficient is set based on the sensitivity of the application scenario to the user's learning ability. The larger the coefficient, the more significant the radius expansion of the area where historical operations are concentrated in the next cycle. A typical value range is 0.3 to 1.0, for example, it can be set to 0.5.
[0036] Spatial two-dimensional kernel density estimation: This is a well-established method in statistics. The input is the two-dimensional coordinates of a sequence of points in the target region. By assigning a kernel function to each point and summing the results over the entire two-dimensional plane, a continuous density distribution surface is obtained. In this step, for any point in the target region, the sum of the "contributions" of all other points within a certain range centered on that point to that point through the kernel function is calculated; the output is the local density value. This value reflects the degree of clustering of historical operation points within the local region centered on that point.
[0037] Normalization: This is a commonly used data processing technique. All acquired local density value sequences are converted into density parameters through linear transformation or other methods. The purpose is to eliminate the influence of absolute numerical magnitude and make density values under different periods and operating habits comparable.
[0038] The statistical characteristics of historical interaction data are used to predict and optimize the scope of future interaction areas. The correction formula for the radius reflects the inherent logic that "the more intensive the operation, the larger the area of focus may be." For example, in map applications, when a user repeatedly zooms and pans at an intersection, their operation points will form a high-density cluster in the intersection area. Based on this, this embodiment automatically expands the interaction focus area of the intersection in the next cycle to ensure that the road information of the entire intersection is clearer, not just the last clicked point.
[0039] Through the above technical solution, this embodiment analyzes the spatial density of historical interaction points, transforming user operating habits into dynamic correction factors for the scope of the interaction focus area. This enables this embodiment to learn the user's refined exploration patterns and adaptively adjust the size of the high-quality encoding area, making the area delineation more closely match the user's actual attention range and improving the accuracy of resource allocation.
[0040] Furthermore, after acquiring the interactive operation fitting polyline within a preset time window containing at least two historical periods, the process further includes: acquiring the coordinate data and timestamp data of all adjacent sampling points of the interactive operation fitting polyline; calculating the displacement based on the coordinate data and combining it with the timestamp data to calculate the instantaneous velocity data; the preset time window contains at least the previous period and the two periods prior; recording the instantaneous velocity data calculated from the last two adjacent sampling points as the terminal instantaneous velocity data; if the terminal instantaneous velocity data of the previous period is higher than a preset high-speed judgment threshold, and the terminal instantaneous velocity data calculated from the two periods prior is lower than a preset low-speed judgment threshold, then inertial compensation processing is performed: acquiring the instantaneous velocity data of the sampling point preceding the sampling point corresponding to the terminal instantaneous velocity data in the previous period; calculating the terminal instantaneous velocity data and the instantaneous velocity data of the previous sampling point. If the instantaneous velocity difference value is less than or equal to a preset end instantaneous velocity threshold, then the end instantaneous velocity data higher than the preset high-speed judgment threshold is used as the prediction reference velocity, multiplied by a preset inertial prediction time constant, to calculate the first inertial slip offset data; the sliding direction vector data is calculated based on the coordinate data of adjacent sampling points at the end of the interactive operation fitted polyline; when delineating the interactive focus area according to the target area point, the coordinate data of each target area point extracted from the interactive operation fitted polyline is translated along the direction of the sliding direction vector data by the first inertial slip offset data to generate a set of inertial compensated target point coordinate data, denoted as the first compensated target point coordinates; the original target area point coordinate data is replaced with the first compensated target point coordinates to obtain the new target area point for the current cycle.
[0041] In this embodiment, a preset high-speed judgment threshold is used to determine whether the user is performing a rapid swipe at a lower speed. This threshold is set based on a speed significantly faster than normal clicking and fine-tuning. A typical value can be between 500 pixels / second and 1000 pixels / second.
[0042] Preset low-speed threshold: The upper limit of the speed used to determine whether the operation has stopped or become extremely slow. The setting is based on a near-stationary operation state. Typical values range from 50 pixels / second to 100 pixels / second.
[0043] The preset instantaneous velocity threshold at the end of the slide is used to determine whether the velocity change at the end of the high-speed slide is a smooth stop or a sharp turn. The setting is based on distinguishing between the acceleration changes caused by inertial deceleration and active braking for steering. Typical values range from 100 pixels / second to 300 pixels / second.
[0044] The preset inertial prediction time constant is a time parameter used to convert velocity into predicted offset. Physically, it represents the time it takes for the content to continue scrolling at its original speed after the user's finger leaves the screen. The setting is based on simulating the inertial scrolling effect common in human-computer interaction, with a typical value ranging from 0.1 to 0.3 seconds.
[0045] Calculate instantaneous velocity data: Calculate the average velocity based on the difference between the coordinate data of adjacent sampling points and the difference between the timestamp data.
[0046] Calculate the sliding direction vector data: Based on the coordinate data of the adjacent sampling points at the end of the fitted polyline by the interactive operation, calculate the vector from the second to last point to the last point, and perform normalization to obtain the unit direction vector.
[0047] This approach identifies the specific interaction pattern of "high-speed swiping followed by a sudden stop" and provides predictive compensation for touch point deviations caused by physiological inertia. The core logic is as follows: when it detects that the previous cycle was still a high-speed swipe while the current cycle has become low-speed, this embodiment predicts that the user's true intention point is not the coordinates of the last touch, but rather the position of a first inertial swipe offset data point extending forward along the swipe direction. This solves the problem of users quickly swiping the map and then suddenly stopping to click, only to have their click misaligned due to finger inertia, allowing the defined interaction focus area to cover the user's true target earlier.
[0048] Through the above technical solution, this embodiment detects specific speed change patterns, predicts the intention point offset caused by operational inertia, and proactively corrects the target area point coordinates. This significantly improves the accuracy of interactive focus area positioning in scenarios where precise operations are performed immediately after rapid browsing, thus enhancing the user experience.
[0049] Furthermore, if the instantaneous velocity difference value is greater than the preset terminal instantaneous velocity threshold, the sampling point corresponding to the instantaneous velocity difference value is recorded as the velocity inflection point. The line segment of the interactive operation fitting polyline corresponding to the previous preset number of sampling points of the velocity inflection point is obtained, and a quadratic polynomial curve is obtained by fitting it using the least squares method. The curvature value of the quadratic polynomial curve at the velocity inflection point is calculated, which is the curvature value of the previous cycle, and is recorded as the first curvature value. The curvature values of the velocity inflection points of the previous two cycles are obtained as the second curvature value. The difference between the first curvature value and the second curvature value is calculated, and the absolute value of the difference is taken to obtain the curvature change. The preset curvature influence attenuation coefficient is multiplied by the curvature change. The first product is obtained; the number 1 is added to the first product to obtain the first sum; the number 1 is divided by the first sum to obtain the attenuation factor; the first inertial sliding offset data is multiplied by the attenuation factor to obtain the second inertial sliding offset data; when delineating the interactive focus area based on the target area point, the coordinate data of each target area point extracted from the fitted polyline of the interactive operation is translated along the direction of the sliding direction vector data by the second inertial sliding offset data to generate a set of inertial compensated target point coordinate data, which is denoted as the second compensated target point coordinates; the second compensated target point coordinates are used to replace the original target area point coordinate data to obtain the new target area point for the current cycle.
[0050] In this embodiment, a preset curvature influence attenuation coefficient is used: this coefficient controls the attenuation strength of curvature changes on inertial offset. The larger the preset curvature influence attenuation coefficient, the stronger the attenuation caused by the same amount of curvature change. The setting is based on the tolerance for uncertainty in orientation adjustment, and the typical value range is 1.0 to 5.0.
[0051] Least squares method for fitting a quadratic polynomial curve: This is a classic curve fitting method. The input is the coordinates of the velocity inflection point and several points preceding it. The quadratic polynomial is solved by minimizing the sum of the squares of the perpendicular distances between the fitted curve and these points.
[0052] Calculate the curvature value: For the fitted quadratic curve, the curvature at a certain point is calculated as a dimensionless scalar, which characterizes the degree of curvature of the curve at that point.
[0053] When this embodiment detects that the user not only stops abruptly but also experiences a drastic change in speed before stopping, potentially indicating a change in direction, more refined inertial compensation is required. Curvature, a geometric feature, is introduced to quantify the "uncertainty of direction adjustment." The physical meaning of the correction formula for the second inertial sliding offset data is: if the curvature of the trajectory at the end of the current abrupt stop operation differs significantly from the previous one, it indicates that the user may be making an adjustment click. The reliability of simple straight-line inertial prediction decreases, thus requiring a significant reduction in the predicted offset. This allows this embodiment to make more conservative predictions when the user performs complex operations such as "rapid sliding to abrupt stop to fine-tuning direction to clicking," avoiding placing high-quality areas in incorrect positions due to overcompensation.
[0054] Through the above technical solution, this embodiment detects the user's potential directional adjustment intention by analyzing the curvature and changes of the sliding end trajectory, and dynamically attenuates the intensity of inertial compensation accordingly. This enhances the robustness of this embodiment in complex, fine-tuning operation scenarios, prevents the inertial compensation mechanism from producing negative effects when directional prediction is inaccurate, and makes the adaptive strategy more intelligent and reliable.
[0055] Furthermore, the process of obtaining the interaction operation fitting polyline is as follows: An initial interaction operation fitting polyline is obtained by performing trajectory extrapolation analysis on the original image interaction operation data; the target device orientation vector is extracted from the original inertial motion data; the target device orientation vector is compared with a preset screen positive orientation vector; if the angle between the target device orientation vector and the preset screen positive orientation vector is less than a preset angle threshold, the initial interaction operation fitting polyline is used as the interaction operation fitting polyline; if the angle between the target device orientation vector and the preset screen positive orientation vector is greater than or equal to the preset angle threshold, a preset rotation mapping relationship table is queried based on the angle between the target device orientation vector and the preset screen positive orientation vector to obtain the corresponding coordinate transformation matrix data; after obtaining the interaction operation fitting polyline, the coordinate data of each point in the initial interaction operation fitting polyline is multiplied by the coordinate transformation matrix data and normalized according to the screen coordinate system to obtain the interaction operation fitting polyline.
[0056] In this embodiment, a preset angle threshold is used to determine whether the device orientation has been actively and significantly rotated by the user. This threshold is set to distinguish between unintentional slight shaking of the device and intentional actions such as switching between portrait and landscape modes. A typical value range is 10 to 20 degrees, for example, it can be set to 15 degrees.
[0057] Trajectory extrapolation analysis: This is a common technique for predicting possible future trajectories from a discrete sequence of touch points, employing methods such as Kalman filtering and multinomial fitting extrapolation. The input is the original image interaction data, and the output is the initial polyline fitted to the interaction.
[0058] Extracting the target device orientation vector from raw inertial motion data: By fusing accelerometer and gyroscope data and using sensor fusion algorithms such as complementary filtering and Kalman filtering, the orientation of the device relative to the Earth coordinate system or the initial coordinate system is calculated, usually represented by quaternions or Euler angles, and can be converted into a three-dimensional orientation vector.
[0059] This implementation addresses a common but overlooked issue of physical characteristic changes during mobile device use: screen rotation. When a user rotates the device, the relationship between the sensor coordinate system and the screen display coordinate system changes abruptly. Without conversion, "left and right swipes" understood based on sensor data may correspond to actual "up and down swipes" on the screen, causing all predictions based on historical trajectories to completely fail.
[0060] This solution compares the target device's orientation vector with a preset screen orientation vector. Upon detecting significant rotation, it applies a rotational coordinate transformation matrix to the coordinates of the initial polygonal line fitted to the interaction operation, transforming it to a new screen coordinate system. This matrix can be quickly retrieved from a pre-calculated mapping table based on the angle between the two vectors. This ensures that regardless of the device's grip orientation, this embodiment's understanding of the user's interaction intent, swipe direction, and target position is always based on the correct screen coordinate system, forming the foundation for the stable operation of the entire adaptive logic in real-world scenarios.
[0061] Through the above technical solution, this embodiment ensures the consistency between the user interaction intent analysis module and the screen display coordinate system by sensing changes in the physical orientation of the device and performing corresponding rotation transformations on the interaction trajectory coordinates. This fundamentally solves the problem of interaction semantic confusion caused by device rotation, guaranteeing the reliability and correctness of the entire adaptive compression transmission method under any device orientation.
[0062] Furthermore, the specific acquisition of the interference image index includes: performing a fast Fourier transform on the original inertial motion data to obtain the frequency domain energy distribution; calculating the total energy exceeding a preset frequency threshold in the frequency domain energy distribution as high-frequency vibration energy; calculating the time domain standard deviation of the original inertial motion data to obtain the time domain jitter amplitude; and weighting and summing the high-frequency vibration energy and the time domain jitter amplitude to obtain the interference image index.
[0063] In this embodiment, a preset frequency threshold is used as the frequency boundary between low-frequency motion and high-frequency vibration. This threshold is set based on the characteristics of human vision in this embodiment; generally, jitter with frequencies higher than 10 to 15 Hz is considered more likely to cause image blurring, while low-frequency jitter causes more image translation. A typical value can be set to 12 Hz.
[0064] The weighting coefficients in the weighted summation are used to balance the contributions of high-frequency vibration energy and temporal jitter amplitude to the final interference image index. For example, each can be assigned a weight of 0.5, indicating that they are equally important; they can also be adjusted empirically, such as assigning a weight of 0.6 to high-frequency vibration and a weight of 0.4 to temporal jitter.
[0065] Fast Fourier Transform (FFT): This transforms raw inertial motion data in the time domain into a frequency domain representation, yielding the frequency domain energy distribution. It is a standard technique in the field of digital signal processing.
[0066] Time-domain standard deviation calculation: The standard deviation is calculated directly on the raw inertial motion data within the time-domain window to obtain the time-domain jitter amplitude, which is used to measure the overall intensity of the motion. This is a basic statistical analysis operation.
[0067] This implementation constructs a comprehensive motion interference quantification index oriented towards image readability. Instead of simply using the raw acceleration amplitude, it extracts features from both the frequency and time domains and then performs weighted fusion. This allows the interference image index to more accurately reflect the actual impact of the current motion state on "image quality," providing a more reliable basis for subsequent compression decisions.
[0068] Through the above technical solution, this embodiment constructs a quantitative index that comprehensively reflects the degree of interference caused by motion to image transmission by integrating motion characteristics in the time and frequency domains. This index provides an objective and accurate basis for this embodiment to determine whether and to what extent adaptive compression is needed, and serves as the data foundation for the entire adaptive logic activation and intensity control.
[0069] Furthermore, the interactive operation fitting polyline is used to obtain several target region points according to a preset segmentation rule. Specifically, this includes: uniformly sampling the interactive operation fitting polyline at preset time intervals to obtain a set of sampling points; processing the set of sampling points using the Douglas-Puk thinning algorithm to retain feature points whose curvature changes exceed a preset curvature threshold to obtain a set of feature points; and using the points in the set of feature points as target region points.
[0070] In this embodiment, the preset time interval is used to uniformly resample the fitted polyline of continuous interactive operations. The setting is based on balancing trajectory reconstruction accuracy and data processing overhead. Too small an interval will retain too many redundant points, increasing computational burden; too large an interval will lose crucial operational details. A typical value range is 50 milliseconds to 200 milliseconds, for example, it can be set to 100 milliseconds.
[0071] Preset curvature threshold: This is a key parameter in the Douglas-Puk thinning algorithm used to determine whether a point should be retained as a feature point. Specifically, it refers to the critical value of the local curvature of the trajectory. The setting is based on the degree of significant change in user intent. When the curvature change of the operation trajectory at a point exceeds this threshold, it indicates that the user may have performed a conscious action such as clicking, turning, or pausing at that point. A typical value needs to be determined experimentally based on the specific application scenario; it is a dimensionless positive number.
[0072] Uniform sampling: This is a fundamental operation in signal processing. Points are selected at equal intervals along the time axis of the interactively fitted polyline, using a preset time interval as the step size, to generate a set of sampling points. If a point on the original polyline is not at the precise sampling time, its coordinates are obtained through linear interpolation.
[0073] Douglas-Puk thinning algorithm: This is a classic linear feature compression algorithm in the field of computer graphics and geographic information systems. Its basic process is as follows: For a curve, retain the start and end points, calculate the perpendicular distance from all intermediate points to the line connecting the start and end points, and find the point with the maximum distance; if this maximum distance is greater than a given threshold, retain the point, and recursively repeat this process for sub-segments; otherwise, discard all intermediate points. In this embodiment, the algorithm is applied to a set of sampling points, and its output is a set of feature points that retain the main shape characteristics of the original trajectory. The "distance tolerance" threshold in the algorithm is geometrically related to the user-defined preset curvature threshold, and together they are used to filter points that can reflect significant changes in trajectory direction.
[0074] This implementation innovatively applies a mature trajectory compression algorithm to the specific stage of "user interaction intent feature extraction." It solves the problem of efficiently and automatically identifying key locations that truly represent the user's intentional actions from a continuous touch stream that may contain noise and redundancy. By first reducing data density through "uniform sampling," and then extracting feature points with significant curvature changes using the "Douglas-Puk algorithm," this combined process effectively filters out data points generated by unconscious micro-jitter during user swiping. This makes the final target area point set more focused on the locations of interactive events such as clicks, long presses, and trajectory inflection points where the user has a clear intention. This provides high-quality, semantically relevant input data for subsequently accurately delineating the interaction focus area.
[0075] Through the above technical solution, this embodiment achieves automatic and accurate extraction of key location points representing the user's core intent from the original operation data by resampling the interaction trajectory and thinning feature points based on curvature changes. This avoids including a large number of irrelevant or redundant touch coordinates in subsequent calculations, significantly improving the data quality and representativeness of the target area points, thus laying a solid foundation for accurate decision-making in the entire adaptive compressed transmission embodiment.
[0076] Figure 3 This is a schematic diagram of the structure of an image semantic adaptive compression transmission system for mobile devices provided in an embodiment of this application. The image semantic adaptive compression transmission system for mobile devices includes: a data acquisition module: used to acquire the original inertial motion data and original image interaction operation data of the target device in the current period, perform time-frequency domain transformation and trajectory extrapolation analysis respectively, and obtain the interference image index and the interaction operation fitting piecewise linear curve; a target region point acquisition module: used to map the interference image index to a global base compression ratio if the interference image index is greater than a preset motion interference threshold, and obtain several target region points from the interaction operation fitting piecewise linear curve according to a preset segmentation rule; an adjustment coefficient acquisition module: used to delineate the interaction focus region of the target region points with a preset radius, obtain several interaction focus regions, and calculate the adjustment coefficient based on the ratio of the intersection area of the several interaction focus regions to the sum of the areas of the interaction focus regions; and a compression parameter acquisition module. The first compression parameter is obtained by reducing the global base compression ratio based on the adjustment coefficient, and the second compression parameter is obtained by increasing the global base compression ratio based on the adjustment coefficient. The second compression parameter is obtained by performing differential encoding on the interactive focus area and the area outside the area according to the first compression parameter and the second compression parameter, respectively, to obtain the first region encoding and the second region encoding, and then merging them into a complete frame bitstream for transmission. The third interaction focus area re-delineation module is obtained by acquiring the interaction operation fitting polyline within a preset time window containing at least two historical periods, extracting the target area point sequence and performing spatial two-dimensional kernel density estimation accordingly, obtaining the local density value of the target area points, and correcting the preset radius accordingly, obtaining the corrected radius, and re-delineating the interactive focus area for the next period.
[0077] This application also provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements an image semantic adaptive compression transmission method for mobile devices.
[0078] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0079] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0080] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0082] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0083] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. An image semantic adaptive compression and transmission method for mobile devices, characterized in that, Includes the following steps: The system acquires the original inertial motion data and original image interaction operation data of the target device in the current cycle. It performs a fast Fourier transform on the original inertial motion data to obtain the frequency domain energy distribution, calculates the high-frequency vibration energy exceeding the preset frequency threshold in the frequency domain, and performs a weighted summation based on the time domain jitter amplitude of the original inertial motion data to obtain the interference image index characterizing the intensity of environmental shaking. It performs trajectory extrapolation analysis on the original image interaction operation data to obtain the initial polyline for the interaction operation fitting, and uses the device orientation vector extracted from the original inertial motion data to query the coordinate transformation matrix data. Based on this, it performs rotation transformation and screen coordinate system normalization on the initial polyline for the interaction operation fitting to obtain the interaction operation fitting polyline. If the interference image index is greater than the preset motion interference threshold, the interference image index is mapped to the global base compression ratio. After the interactive operation fitted polyline is uniformly sampled at a preset time interval, the feature points whose curvature changes exceed the preset curvature threshold are thinned out and retained as target region points. The difference between the interference image index and the preset motion interference threshold is calculated. Multiply the difference value by a preset radius scaling factor to obtain a radius adjustment amount; add the radius adjustment amount to a preset base radius value, and the result is the preset radius used to define the interactive focus area; The interactive focus area of the target area is defined by a preset radius, resulting in several interactive focus areas. The ratio of the intersection area of these interactive focus areas to the sum of the areas of the interactive focus areas is recorded as a first ratio. The intersection areas of interactive focus areas whose first ratio is less than a first preset ratio are counted as a first quantity, and the intersection areas of interactive focus areas whose first ratio is greater than or equal to the first preset ratio are counted as a second quantity. When the first quantity is greater than or equal to the second quantity, a first preset coefficient is multiplied by a first preset magnification factor to obtain an adjustment coefficient. When the second quantity is greater than the first quantity, the first preset coefficient is multiplied by a first preset reduction factor to obtain an adjustment coefficient. The first compression parameter is obtained by decreasing the global base compression ratio based on the adjustment factor, and the second compression parameter is obtained by increasing the global base compression ratio based on the adjustment factor. The interactive focus area is encoded using a first compression parameter to obtain a first region code, and the area outside the interactive focus area is encoded using a second compression parameter to obtain a second region code, and then merged into a complete frame bitstream for transmission. The system obtains a fitted polyline of interactive operations within a preset time window containing at least two historical periods. It extracts the target area point sequence and performs spatial two-dimensional kernel density estimation based on it to obtain the local density value of the target area point. The local density value of the target area point is normalized to obtain the density parameter. When the density parameter is greater than the preset density threshold, a correction factor is generated by combining the preset density influence coefficient. The preset radius is multiplied by the correction factor to obtain the correction radius. The interactive focus area is then redefined based on this and used for the next period.
2. The image semantic adaptive compression and transmission method for mobile devices according to claim 1, characterized in that, The specific process for obtaining the adjustment coefficient is as follows: The ratio of the area of the intersection of several interactive focal areas to the sum of the areas of the interactive focal areas is denoted as the first ratio. The first quantity is the intersection area of interactive focus areas where the first ratio is less than the first preset ratio. The second quantity is the intersection area of interactive focus areas where the first ratio is greater than or equal to the first preset ratio. When the first quantity is greater than or equal to the second quantity, it is determined that the interactive focus area is dispersed, and the first preset coefficient is multiplied by the first preset amplification factor to obtain the adjustment coefficient. When the second quantity is greater than the first quantity, it is determined that the interactive focus area is overlapping. The first preset coefficient is multiplied by the first preset reduction factor to obtain the adjustment coefficient.
3. The image semantic adaptive compression and transmission method for mobile devices according to claim 1, characterized in that, The specific process for obtaining the correction radius is as follows: Taking any target region point in the target region point sequence as the center, a circular neighborhood is defined with a preset radius. The weighted contribution of other target region points in the circular neighborhood to the target region point is processed by spatial two-dimensional kernel density estimation to obtain the local density value of the target region point. The local density values of all target region points are normalized to obtain the density parameters of the target region points; If the density parameter of the target area points is less than or equal to the preset density threshold, the preset radius will not be corrected. If the density parameter of the target area point is greater than the preset density threshold, the correction factor is obtained by adding the product of the preset density influence coefficient and the density parameter to the number 1. The preset radius is then multiplied by the correction factor to obtain the corrected radius of the target area point.
4. The image semantic adaptive compression and transmission method for mobile devices according to claim 1, characterized in that, After obtaining the fitted line of interactive operations within a preset time window containing at least two historical periods, the method further includes: Obtain the coordinate data and timestamp data of all adjacent sampling points of the fitted polyline in the interactive operation; The displacement is calculated based on the coordinate data, and the instantaneous velocity data is calculated by combining the timestamp data. The preset time window must include at least the previous cycle and the two previous cycles; The instantaneous velocity data calculated from the last two adjacent sampling points is recorded as the terminal instantaneous velocity data; If the terminal instantaneous velocity data of the previous cycle is higher than the preset high-speed determination threshold, and the terminal instantaneous velocity data calculated in the previous two cycles is lower than the preset low-speed determination threshold, then inertial compensation processing is performed: Obtain the instantaneous velocity data of the sampling point preceding the sampling point corresponding to the instantaneous velocity data at the end of the previous cycle; The instantaneous velocity difference between the terminal instantaneous velocity data and the instantaneous velocity data of the previous sampling point is calculated. If the instantaneous velocity difference is less than or equal to a preset terminal instantaneous velocity threshold, the terminal instantaneous velocity data that is higher than the preset high-speed determination threshold is used as the prediction reference velocity, multiplied by a preset inertial prediction time constant, to calculate the first inertial slip offset data. The sliding direction vector data is calculated based on the coordinate data of adjacent sampling points at the end of the fitted polyline using the interactive operation. When defining the interactive focus area based on the target area point, the coordinate data of each target area point extracted from the fitted polyline of the interactive operation is translated along the direction of the sliding direction vector data by the first inertial sliding offset data to generate a set of inertial compensated target point coordinate data, which is denoted as the first compensated target point coordinates. The coordinates of the first compensation target point are used to replace the original coordinate data of the target area point to obtain the new target area point for the current period.
5. The image semantic adaptive compression and transmission method for mobile devices according to claim 4, characterized in that, If the instantaneous velocity difference value is greater than the preset terminal instantaneous velocity threshold, the sampling point corresponding to the instantaneous velocity difference value is recorded as the velocity inflection point. The line segment of the interactive operation fitting polyline corresponding to the preset number of sampling points of the velocity inflection point is obtained, and a quadratic polynomial curve is obtained by fitting it using the least squares method. The curvature value of the quadratic polynomial curve at the velocity inflection point is calculated, which is the curvature value of the previous cycle, and is denoted as the first curvature value. Obtain the curvature value of the velocity inflection point of the previous two cycles, and use it as the second curvature value; Calculate the difference between the first curvature value and the second curvature value, and take the absolute value of the difference to obtain the curvature change. Multiply the curvature change by the preset curvature influence attenuation coefficient to obtain the first product. Add the number one to the first product to obtain the first sum; Divide the number one by the first sum to obtain the attenuation factor; Multiply the first inertial slip offset data by the attenuation factor to obtain the second inertial slip offset data; When defining the interactive focus area based on the target area point, the coordinate data of each target area point extracted from the fitted polyline of the interactive operation is translated along the direction of the sliding direction vector data by the second inertial sliding offset data to generate a set of inertial compensated target point coordinate data, which is denoted as the second compensated target point coordinates. The coordinates of the second compensation target point are used to replace the original coordinate data of the target area point to obtain the new target area point for the current period.
6. The image semantic adaptive compression and transmission method for mobile devices according to claim 1, characterized in that, The process of obtaining the fitted polyline in the interactive operation is as follows: By performing trajectory extrapolation analysis on the original image interaction operation data, an initial polyline for fitting the interaction operation is obtained; Extract the target device orientation vector from the raw inertial motion data; The target device orientation vector is compared with a preset screen positive orientation vector; If the angle between the target device orientation vector and the preset screen positive orientation vector is less than a preset angle threshold, then the initial polyline of the interaction operation fitting will be used as the interaction operation fitting polyline. If the angle between the target device orientation vector and the preset screen positive orientation vector is greater than or equal to a preset angle threshold, then the preset rotation mapping relationship table is queried according to the angle between the target device orientation vector and the preset screen positive orientation vector to obtain the corresponding coordinate transformation matrix data. After obtaining the interaction operation fitting polyline, the coordinate data of each point in the initial interaction operation fitting polyline is multiplied with the coordinate transformation matrix data and normalized according to the screen coordinate system to obtain the interaction operation fitting polyline.
7. The image semantic adaptive compression and transmission method for mobile devices according to claim 1, characterized in that, The specific acquisition of the interference image index includes: The original inertial motion data is subjected to a fast Fourier transform to obtain the frequency domain energy distribution; The total energy exceeding a preset frequency threshold in the frequency domain energy distribution is calculated as the high-frequency vibration energy; The time-domain standard deviation of the original inertial motion data is calculated to obtain the time-domain jitter amplitude; The interference image index is obtained by weighting and summing the high-frequency vibration energy and the time-domain jitter amplitude.
8. The image semantic adaptive compression and transmission method for mobile devices according to claim 1, characterized in that, The interactive operation fits a polyline and obtains several target region points according to a preset segmentation rule, specifically including: The interactive operation fitting polyline is uniformly sampled at preset time intervals to obtain a set of sampling points; The sampling point set is processed by the Douglas-Puk thinning algorithm to retain feature points whose curvature changes exceed a preset curvature threshold, thus obtaining a feature point set. The points in the set of feature points are taken as the target region points.
9. An image semantic adaptive compression and transmission system for mobile devices, characterized in that, include: Data acquisition module: used to acquire the original inertial motion data and original image interaction operation data of the target device in the current cycle, perform fast Fourier transform on the original inertial motion data to obtain the frequency domain energy distribution, calculate the high-frequency vibration energy exceeding the preset frequency threshold in the frequency domain, and perform weighted summation with the time domain jitter amplitude of the original inertial motion data to obtain the interference image index characterizing the intensity of environmental shaking, perform trajectory extrapolation analysis on the original image interaction operation data to obtain the initial polyline of interaction operation fitting, and use the device orientation vector extracted from the original inertial motion data to query the coordinate transformation matrix data, and perform rotation transformation and screen coordinate system normalization on the initial polyline of interaction operation fitting accordingly to obtain the interaction operation fitting polyline; Target region point acquisition module: If the interference image index is greater than the preset motion interference threshold, the interference image index is mapped to the global base compression ratio, the interactive operation fitted polyline is uniformly sampled at preset time intervals, and feature points whose curvature changes exceed the preset curvature threshold are thinned out and retained as target region points, and the difference between the interference image index and the preset motion interference threshold is calculated. Multiply the difference value by a preset radius scaling factor to obtain a radius adjustment amount; add the radius adjustment amount to a preset base radius value, and the result is the preset radius used to define the interactive focus area; The calibration coefficient acquisition module is used to delineate the interactive focus area of the target area point with a preset radius, obtain several interactive focus areas, and record the ratio of the intersection area of the several interactive focus areas to the sum of the areas of the interactive focus areas as a first ratio. The number of intersection areas of interactive focus areas whose first ratio is less than a first preset ratio is counted as a first quantity, and the number of intersection areas of interactive focus areas whose first ratio is greater than or equal to the first preset ratio is counted as a second quantity. When the first quantity is greater than or equal to the second quantity, the first preset coefficient is multiplied by a first preset magnification factor to obtain the calibration coefficient. When the second quantity is greater than the first quantity, the first preset coefficient is multiplied by a first preset reduction factor to obtain the calibration coefficient. Compression parameter acquisition module: used to decrease the global base compression ratio according to the adjustment coefficient to obtain the first compression parameter, and increase the global base compression ratio according to the adjustment coefficient to obtain the second compression parameter; Image encoding and transmission module: used to encode the interactive focus area using a first compression parameter to obtain a first region encoding, and to encode the area outside the interactive focus area using a second compression parameter to obtain a second region encoding, and to merge them into a complete frame bitstream for transmission. The interactive focus area re-delineation module is used to obtain the interactive operation fitting polyline within a preset time window containing at least two historical periods, extract the target area point sequence and perform spatial two-dimensional kernel density estimation based on it to obtain the local density value of the target area point, normalize the local density value of the target area point to obtain the density parameter, when the density parameter is greater than the preset density threshold, combine the preset density influence coefficient to generate a correction factor, and multiply the preset radius by the correction factor to obtain the correction radius, and redefine the interactive focus area based on it for the next period.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 8.