Interactive processing method and system for digital media file
By detecting the user's three-dimensional spatial position and orientation to generate personalized display parameters, the problem of traditional digital media systems being unable to adapt to user needs is solved, achieving a highly immersive and naturally interactive digital media experience.
Patent Information
- Application Number
- CN202512035755.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2045-12-31
AI Technical Summary
Traditional digital media interactive systems cannot be personalized to meet the individual viewing needs and comfort levels of users, resulting in a poor user experience.
By detecting the target user's coordinate position, visual height, and visual orientation angle in the digital interaction area, personalized display parameters are generated, including the projection center point, projection size, and projection angle. Based on these parameters, content is selected from the media library for processing, responding to the user's gesture interaction actions.
It significantly improves the personalization level of digital media interaction systems and the immersive experience of users, solves the problems of viewing angle deviation and stiff interaction caused by fixed projection mode, and achieves a natural and smooth interactive experience.
Smart Images

Figure CN121433554A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of human-computer interaction, and in particular to an interactive processing method and system of digital media files. BACKGROUND
[0002] With the rapid development of digital technology, digital media files are increasingly widely used in fields such as exhibition halls, museums, and educational institutions, providing users with rich visual experiences and interactive opportunities. Traditional digital media interactive systems usually rely on fixed display devices and predetermined display content. For example, in some digital art exhibitions, digital art is presented to all users in a uniform manner, but this display method ignores the viewing needs and comfort of individual users, resulting in users not being able to obtain the best viewing experience and poor digital media interactive experience. SUMMARY
[0003] The present application provides an interactive processing method and system of digital media files to solve the technical problem of poor user interaction experience caused by the lack of personalized adaptation of digital media files in the prior art.
[0004] In view of the above problems, the present application provides an interactive processing method and system of digital media files.
[0005] In a first aspect, the present application provides an interactive processing method of digital media files, comprising: detecting user positioning information of a target user entering a digital interactive area, the user positioning information including the coordinate position, visual height, and visual orientation angle of the target user in the digital interactive area; generating personalized display parameters for the target user according to the coordinate position, visual height, and visual orientation angle, the personalized display parameters including a projection center point, a projection size, and a projection angle; selecting corresponding media content in a digital media file library based on the projection center point, projection size, and projection angle to process and obtain initial projection content, the initial projection content including a plurality of digital media objects; in response to a gesture interaction action of the target user, performing interactive processing based on the gesture interaction action and the initial projection content to generate interactive response content.
[0006] In a second aspect, the present application provides an interactive processing system of digital media files, comprising: a user positioning acquisition module for detecting user positioning information of a target user entering a digital interactive area, the user positioning information including the coordinate position, visual height, and visual orientation angle of the target user in the digital interactive area; The display parameter generation module is configured to generate personalized display parameters for the target user according to the coordinate position, the visual height and the visual orientation angle, wherein the personalized display parameters include a projection center point, a projection size and a projection angle. The projection content acquisition module is configured to process corresponding media content in a digital media file library based on the projection center point, the projection size and the projection angle to obtain initial projection content, wherein the initial projection content includes a plurality of digital media objects. The interaction response module is configured to respond to a gesture interaction action of the target user, perform interaction processing based on the gesture interaction action and the initial projection content, and generate interaction response content.
[0007] One or more technical solutions provided in the present application have at least the following technical effects or advantages: The present application provides an interactive processing method and system for digital media files. The coordinate position, visual height and visual orientation angle of a target user in a digital interaction area are detected in real time, and personalized display parameters including a projection center point, a projection size and a projection angle are dynamically generated based on the above parameters. Then, the initial projection is formed by intelligently selecting and processing the content from the media library based on the above parameters. Finally, the corresponding interaction response content is generated in response to the natural gesture interaction action of the user. The personalized level of the digital media interaction system, the immersion of the user experience and the natural fluency of the interaction are significantly improved. Compared with the traditional method, the technical solution provided by the present application significantly overcomes the problems of visual angle deviation, content distortion and interaction stiffness caused by the fixed projection mode. The technical effect of dynamically sensing the user state, intelligently adapting the display content and accurately responding to the natural interaction intention of the digital media interaction system is achieved. The highly personalized, immersive and smooth digital media interaction experience is provided for the user. BRIEF DESCRIPTION OF DRAWINGS
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0009] Figure 1 A flowchart of an interactive processing method for digital media files provided by an embodiment of the present application.
[0010] Figure 2 A structural diagram of an interactive processing system for digital media files provided by an embodiment of the present application.
[0011] In the drawings, the components represented by the numbers are described as follows: User location acquisition module 100, display parameter generation module 200, projection content acquisition module 300, and interactive response module 400. Detailed Implementation
[0012] This application provides an interactive processing method and system for digital media files, which addresses the technical problem of poor user interaction experience caused by the lack of personalized adaptation of digital media files in the prior art.
[0013] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0014] It should be noted that the terms "comprising" and "having" are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or server that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or modules that are not explicitly listed or that are inherent to these processes, methods, products, or devices.
[0015] Example 1, as Figure 1 As shown, this application provides an interactive processing method for digital media files, wherein the method includes: S10: Detect the user positioning information of the target user entering the digital interaction area. The user positioning information includes the target user's coordinate position, visual height, and visual orientation angle in the digital interaction area.
[0016] Traditional digital media display methods can only detect the user's two-dimensional position, but cannot obtain key three-dimensional spatial information such as the user's visual height and head orientation. This makes it impossible for the system to truly understand the user's actual viewing angle, resulting in a lack of accurate data basis for subsequent projection adjustments. The final content may deviate from the user's optimal viewing angle, resulting in a poor viewing experience.
[0017] Step S10 in the method provided in this application embodiment includes: A three-dimensional coordinate system for the digital interactive area is established with the central projection device of the digital interactive area as the center point, the preset response distance as the radius, and the preset vertical detection range as the height. The position information of the target user in the three-dimensional coordinate system is detected to obtain the coordinate position of the target user in the digital interaction area; Detect the head height position of the target user, and determine the visual height of the target user based on the head height position; detect a head vertical orientation of the target user, and determine a visual orientation angle of the target user according to the head vertical orientation; take the coordinate position of the target user in the digital interaction area, the visual height, and the visual orientation angle as the user positioning information.
[0018] In the embodiments of the present application, the user positioning information of the target user entering the digital interaction area is detected, and the user positioning information includes the coordinate position, the visual height, and the visual orientation angle of the target user in the digital interaction area.
[0019] Specifically, first, a central projection device of the digital interaction area is taken as a center point (0, 0, 0) of a three-dimensional coordinate system, a preset response distance is taken as a radius, and a preset vertical detection range is taken as a height, to establish a three-dimensional coordinate system of the digital interaction area. The preset response distance is a horizontal range in which the central projection device can respond to user interaction, which is obtained based on the response performance of the central projection device itself, for example, 5 meters. The preset vertical detection range is a horizontal distance in which the central projection device can respond to user interaction, which is obtained based on the response performance of the central projection device itself, for example, 2.5 meters.
[0020] Further, the position information of the target user in the three-dimensional coordinate system is detected to obtain the coordinate position of the target user in the digital interaction area. For example, the position of the torso of the target user in space is captured in real time by a space tracking locator. The position is mapped with the established three-dimensional coordinate system to calculate the X-axis and Y-axis coordinates of the current position of the user relative to the origin of the coordinate system, for example, a coordinate value such as (1.3, 1.5) is obtained as the coordinate position.
[0021] Further, the head height position of the target user is detected, and the visual height of the target user is determined according to the head height position. For example, the key node of the head of the target user, i.e., the position between the eyes, is identified by image detection through an installed camera, and the Y-axis coordinate value of the node in the three-dimensional coordinate system is obtained. The Y-axis coordinate value of the head node obtained is directly determined as the visual height of the target user. For example, if the Y coordinate of the head node of the user is detected to be 1.65 meters, then the visual height of the target user is 1.65 meters.
[0022] Further, the head vertical orientation of the target user is detected, and the visual orientation angle of the target user is determined according to the head vertical orientation. For example, the inclination angle of the head of the user relative to the horizontal plane is obtained by using an inclination angle sensor. The horizontal forward direction is taken as the 0-degree line of sight reference, the head of the user is raised upward as the positive angle, and the head of the user is lowered downward as the negative angle. For example, if the head of the user is raised 30 degrees relative to the horizontal plane, then the visual orientation angle of the target user is 30 degrees.
[0023] The coordinate position, visual height and visual orientation angle of the target user in the digital interaction area are taken as user positioning information, for example, one user positioning information is (1.3, 1.5, 1.65, 30°), which represents that the user's torso is at (1.3, 1.5), the head height position is 1.65 meters, and the visual orientation angle is 30 degrees.
[0024] By establishing a three-dimensional coordinate system of the digital interaction area and accurately detecting the coordinate position, visual height and visual orientation angle of the target user in the coordinate system, the core technical effect of providing high-precision, multi-dimensional user space state data for the entire interaction system is achieved. The relative plane relationship between the user and the projection device is determined by obtaining the coordinate position; the vertical viewing level of the user is determined by detecting the visual height, ensuring that the content is not too high or too low; and the angle between the user's line of sight and the projection plane is determined by identifying the visual orientation angle, laying a reliable data foundation for subsequent generation of truly personalized display parameters.
[0025] S20: generating personalized display parameters for the target user according to the coordinate position, the visual height and the visual orientation angle, the personalized display parameters including a projection center point, a projection size and a projection angle.
[0026] After obtaining the user positioning information, how to convert these space data into personalized display parameters that can guide content presentation is a technical problem that needs to be solved. If the projection center point cannot be obtained according to the user's position and viewing angle, the core area of the content may not coincide with the user's visual focus; if the projection size cannot be dynamically adjusted according to the viewing distance, the user may see images with distorted proportions or unclear details; and if the projection angle cannot be compensated according to the visual orientation, visual deviation similar to image tilt or trapezoidal distortion will occur.
[0027] The step S20 in the method provided by the embodiment of the application includes: connecting the coordinate position of the target user and the position of the central projection device to determine the direction of the connecting line and the user viewing distance; extending the preset viewing distance from the coordinate position of the target user along the direction of the connecting line to the central projection device to obtain a horizontal viewing point; determining a vertical viewing point according to the visual height, and determining a reference projection point based on the horizontal viewing point and the vertical viewing point; correcting the reference projection point by an arc line based on the visual orientation angle to obtain a projection center point; performing size lookup in a distance size mapping table according to the user viewing distance to obtain a projection size; determining a projection angle based on the visual orientation angle, wherein the projection angle and the visual orientation angle are in a complementary relationship; The projection center point, the projection size and the projection angle are taken as the personalized display parameters of the target user.
[0028] In the embodiments of the present application, the personalized display parameters for the target user are generated according to the coordinate position, the visual height and the visual orientation angle, and the personalized display parameters include the projection center point, the projection size and the projection angle.
[0029] Specifically, first, the coordinate position of the target user is connected with the position of the central projection device to determine the connection line direction and the user viewing distance. For example, the coordinate position of the target user is (1.3, 1.5), and the position of the central projection device is (0, 0), then the vector (-1.3, -1.5) from the user position to the central projection device position is obtained as the connection line direction. The modulus of the vector from the user position to the central projection device position is calculated as the user viewing distance, for example, the modulus of (-1.3, -1.5) is 1.98, and the user viewing distance is 1.98.
[0030] Further, the horizontal viewing point is obtained by extending the preset viewing distance from the coordinate position of the target user along the connection line direction to the central projection device. Specifically, the preset viewing distance is a preset distance that can maximize the user viewing experience, and is set to 1.5 meters for example. The horizontal viewing point is obtained by extending the preset viewing distance from the coordinate position of the target user along the connection line direction to the central projection device, for example, the connection line direction is (-1.3, -1.5), and the preset viewing distance is 1.5 meters, then the horizontal viewing point is (0.32, 0.37).
[0031] Further, the vertical viewing point is determined according to the visual height, and the reference projection point is determined based on the horizontal viewing point and the vertical viewing point. For example, the visual height is determined as the vertical viewing point, for example, the visual height of the target user is 1.65, then 1.65 is determined as the vertical viewing point. The reference projection point is determined based on the horizontal viewing point and the vertical viewing point, for example, the horizontal viewing point is (0.32, 0.37), and the vertical viewing point is 1.65, then the reference projection point is (0.32, 0.37, 1.65).
[0032] Further, the reference projection point is arc corrected based on the visual orientation angle to obtain a projection center point. Specifically, taking the central projection device as the center and the preset viewing distance as the radius, the reference projection point is adjusted along an arc trajectory in the vertical direction according to the visual orientation angle: when the visual orientation angle is positive, that is, looking up, the reference projection point is adjusted upward along the arc; when the visual orientation angle is negative, that is, looking down, the reference projection point is adjusted downward. The vertical offset of the adjustment can be calculated by the formula: vertical offset = preset viewing distance × sin (visual orientation angle). For example, when the visual orientation angle is 30 degrees and the preset viewing distance is 1.5 meters, the vertical offset = 1.5 × sin (30°) = 0.75 meters. The Z coordinate of the reference projection point is added by the vertical offset to obtain the corrected Z coordinate. For example, when the Z coordinate of the reference projection point is 1.65 meters, the corrected Z coordinate = 1.65 + 0.75 = 2.4 meters. The corrected projection center point is (0.32, 0.37, 2.4). The projection center point after arc correction is more consistent with the arc trajectory of the natural line of sight when the user looks up or looks down, and has the effect of improving the user's viewing experience.
[0033] According to the user viewing distance, the size is looked up in the distance-size mapping table to obtain a projection size. The distance-size mapping table is a mapping table that records the viewing distance and the corresponding optimal projection size, and is set in advance. For example, when the user viewing distance is 1.98, the size is looked up in the distance-size mapping table to obtain a projection size of 4 meters × 3 meters.
[0034] The projection angle is determined based on the visual orientation angle, wherein the projection angle and the visual orientation angle are in a complementary relationship. For example, according to the rule that the projection angle and the visual orientation angle are in a complementary relationship, that is, the values are equal and the directions are opposite, the projection angle is calculated. For example, when the user visual orientation angle is 30 degrees upward, the projection angle is 30 degrees downward, that is, -30 degrees. The obtained projection angle can ensure that the projection surface is approximately perpendicular to the user's line of sight, so that the user can see the front effect without geometric distortion.
[0035] The projection center point, the projection size and the projection angle are taken as the personalized display parameters of the target user. For example, the personalized display parameters of the target user can be the projection center point (0.32, 0.37, 2.4), the projection size 4 meters × 3 meters, and the projection angle -30 degrees.
[0036] The coordinate position, the visual height and the visual orientation angle are comprehensively processed through a series of algorithms, personalized display parameters including a projection center point, a projection size and a projection angle are generated, and a technical effect of converting user space posture data into specific and executable display instructions is achieved. By determining the connection line direction, the user viewing distance and performing arc correction, the projection center point can ensure that the core visual area of the content is accurately aligned with the user's line of sight; by mapping the projection size based on the viewing distance, it is ensured that the display content can be presented in the most suitable visual proportion for the current distance regardless of the distance of the user; by establishing a complementary relationship between the projection angle and the visual orientation angle, the image geometric distortion that may be caused by the user's visual angle deflection is effectively compensated, and the visual orthodoxy of the display content is ensured.
[0037] S30: based on the projection center point, the projection size and the projection angle, selecting corresponding media content in the digital media file library for processing to obtain initial projection content, the initial projection content including a plurality of digital media objects.
[0038] The calling and layout of the content library of the traditional digital media system are often fixed, and cannot be dynamically adapted according to the real-time changing projection size, projection center point and other parameters. This may lead to the problem that even if the projection parameters have been optimized for the user, the displayed content itself may be sparse or crowded in the new projection parameters due to the rigid layout, key elements may be truncated or deviated from the center, or the content resolution and the projection size do not match, resulting in blurring and other problems.
[0039] The method provided in the embodiments of the present application includes step S30: determining a preset grid layout according to the projection size, and determining the number requirement of the digital media objects based on the preset grid layout; randomly selecting a plurality of digital media objects from the digital media file library based on the number requirement, the digital media objects including three-dimensional media models and plane media images; obtaining first resolution data of each digital media object as initial display data of the plurality of digital media objects; obtaining basic projection content based on the preset grid layout and the initial display data of the plurality of digital media objects; adjusting the position and the angle of the basic projection content based on the projection center point and the projection angle to obtain the initial projection content.
[0040] In the embodiments of the present application, based on the projection center point, the projection size and the projection angle, corresponding media content in the digital media file library is selected for processing to obtain initial projection content, and the initial projection content includes a plurality of digital media objects.
[0041] Specifically, first, a preset grid layout is determined according to the projection size, and the quantity requirement of digital media is determined based on the preset grid layout. For example, the projection size is 4m x 3m, and the preset grid layout is obtained as 4 x 3, i.e., a 3-row 4-column grid layout, by looking up in a mapping relationship table of projection size-grid layout. The mapping relationship table of projection size-grid layout is pre-set and defines the recommended grid layout corresponding to different projection size ranges. Based on the preset grid layout, the quantity requirement of digital media is determined, for example, 4 x 3 grid layout requires 4 x 3 = 12 digital media contents. For example, the digital media content can be a three-dimensional artistic model or a flat image work.
[0042] Further, based on the quantity requirement, a plurality of digital media objects are randomly selected from a digital media file library. The digital media objects include three-dimensional media models such as sculptures or object models in obj or fbx format, and flat media images such as paintings or photos in jpg or png format. Using a random number generator, 12 digital media objects are randomly selected from the digital media file library without repetition.
[0043] Further, first resolution data of each digital media object is obtained as initial display data of the plurality of digital media objects. When each digital media object is stored in the library, 3 resolutions of display data are pre-generated and associatedly stored, such as high resolution data, medium resolution data, and low resolution data. The first resolution data can be a medium resolution version. The 12 randomly selected digital media objects are traversed, and their preset first resolution data is directly read from the associated data file. For example, for a three-dimensional media model, the first resolution data can be a low-precision model containing 2500 polygons; for a flat media image, the first resolution data can be an 800 x 600 pixel image.
[0044] Further, based on the preset grid layout and the initial display data of the plurality of digital media objects, a basic projection content is obtained. Specifically, based on the determined preset grid layout such as 4 x 3 grid, the randomly selected plurality of digital media objects such as 12 objects, and the initial display data corresponding to these objects such as 5 low-precision models and 7 800 x 600 resolution images, each digital media object is placed in a cell of the grid according to the grid layout with its first resolution data to form a complete projection picture as the basic projection content.
[0045] Further, based on the projection center point and the projection angle, the position and angle of the basic projection content are adjusted to obtain initial projection content. For example, the basic projection content is translated to the projection center point (0.32, 0.37, 2.4) and adjusted according to the projection angle, such as being inclined downward by 30 degrees, to obtain the initial projection content optimized and calibrated according to the personalized display parameters of the target user.
[0046] By determining the grid layout and the content quantity requirement according to the projection size, and adjusting the position and angle of the selected content based on the projection center point and the projection angle, the technical effect of making the media content itself and its initial presentation state highly adaptive to the personalized display environment is achieved. By dynamically selecting a proper number of digital media objects from the media library and preliminarily laying out the selected digital media objects according to the optimized grid, the basic projection content is formed. Further, by adjusting the projection center point and the angle, it is ensured that the digital media content can be accurately placed on the best display position calculated for the user, and it is ensured that the final initial projection content can perfectly adapt to the viewing position and angle of the user, thereby laying a good foundation for subsequent interactive operations.
[0047] S40: in response to the gesture interaction action of the target user, performing interactive processing based on the gesture interaction action and the initial projection content to generate interactive response content.
[0048] The prior art generally has the problem of single response mode in the interactive level, which seriously affects the depth and fluency of user experience. For example, relying on physical controllers or simple predefined actions, the natural and diversified gesture intentions of the user cannot be accurately recognized and understood, resulting in the user often encountering the dilemma of interactive delay, misrecognition or limited functions when deeply interacting with the content of interest.
[0049] The step S40 in the method provided by the embodiments of the present application includes: detecting a gesture interaction action of the target user, the gesture interaction action including a sliding gesture, a selection gesture and a zoom-in gesture; when the sliding gesture is detected, identifying a direction of the sliding gesture, re-determining a plurality of updated digital media objects from the digital media file library according to the direction of the sliding gesture, obtaining first resolution data of the plurality of updated digital media objects, and generating interactive response content; when the selection gesture is detected, identifying a digital media object pointed by the selection gesture, determining a target digital media object, obtaining second resolution data of the target digital media object, and generating interactive response content; wherein the identifying the digital media object pointed by the selection gesture, the determining the target digital media object, the obtaining the second resolution data of the target digital media object, and the generating the interactive response content include: detecting a pointing position of the selection gesture in the projected content; determining a corresponding grid cell according to the pointing position, and identifying a digital media object in the grid cell, and determining the identified digital media object as the target digital media object; obtaining second resolution data of the target digital media object from the digital media file library; processing the target digital media object based on the second resolution data, and switching the target digital media object to a separate display mode; displaying the target digital media object in the separate display mode as the interactive response content; when a zoom-in gesture is detected, identifying a local area in the target digital media object pointed by the zoom-in gesture, determining a target local area, obtaining third resolution data of the target local area, and generating interactive response content; wherein the identifying of the local area in the target digital media object pointed by the zoom-in gesture, the determining of the target local area, the obtaining of the third resolution data of the target local area, and the generating of the interactive response content, comprise: detecting a pointing coordinate and a gesture amplitude of the zoom-in gesture; determining a corresponding local area center position in the target digital media object according to the pointing coordinate, and determining a zoom-in multiple of the local area based on the gesture amplitude; determining the target local area based on the center position, the zoom-in multiple, and the projection size; obtaining third resolution data of the target digital media object from the digital media file library; cutting the third resolution data of the target digital media object according to the target local area, to obtain third resolution data of the target local area; processing the target local area for zoom-in display based on the third resolution data of the target local area, and switching the target local area to a zoom-in display mode; displaying the target local area in the zoom-in display mode as the interactive response content.
[0050] In the embodiments of the present application, in response to the gesture interaction action of the target user, interactive processing is performed based on the gesture interaction action and the initial projected content, and interactive response content is generated.
[0051] Specifically, first, a gesture interaction action of a target user is detected, wherein the gesture interaction action includes a sliding gesture, a selection gesture, and a zoom-in gesture. Illustratively, a gesture recognition sensor is used to capture the motion trajectory and posture of the user's hand in real time, and a preset gesture recognition algorithm such as a skeleton tracking algorithm is used to analyze and classify the gesture type: the sliding gesture is defined as the continuous movement of the hand in the horizontal or vertical direction; the selection gesture is defined as the short stay pointing action of the hand; and the zoom-in gesture is defined as the action of separating or pinching the hands.
[0052] Further, when the sliding gesture is detected, the direction of the sliding gesture is recognized, a plurality of updated digital media objects are re-determined from the digital media file library according to the direction of the sliding gesture, first resolution data of the plurality of updated digital media objects is acquired, and an interactive response content is generated. Illustratively, it is detected that the gesture of the target user is a right sliding, a plurality of updated digital media objects are randomly determined from the digital media file library, such as 12 updated digital media objects are re-determined according to the grid layout, first resolution data of the plurality of updated digital media objects is acquired, and an interactive response content is generated, such as the user gesture is a right sliding, the digital media objects are transformed into updated digital media objects from left to right.
[0053] When the selection gesture is detected, the digital media object pointed by the selection gesture is recognized, the target digital media object is determined, the second resolution data of the target digital media object is acquired, and an interactive response content is generated.
[0054] Specifically, based on the gesture recognition sensor, the pointing position of the selection gesture in the projection content is detected. A coordinate system is established with the lower left corner of the projection content as the origin, the horizontal length as the X-axis, and the vertical width as the Y-axis, and the pointing position is the coordinate point pointed by the gesture in the projection content.
[0055] Further, according to the pointing position, the corresponding grid cell is determined, i.e., the grid cell in which the pointing position coordinate point is located is determined, and the digital media object in the grid cell is recognized, and the recognized digital media object such as a digital media image of a first resolution is determined as the target digital media object.
[0056] From the digital media file library, the second resolution data of the target digital media object is acquired. For example, the first resolution of the target digital media image is 800x600, and the second resolution data is 2400x1800.
[0057] Based on the second resolution data, the target digital media object is processed, and the target digital media object is switched to a separate display mode, i.e., other projection content is canceled, and only the target digital media object displayed with the second resolution data is retained.
[0058] The target digital media object displayed separately is taken as the interactive response content in response to the selection gesture of the user.
[0059] When the zoom-in gesture is detected, a local area in the target digital media object pointed by the zoom-in gesture is identified, the target local area is determined, third resolution data of the target local area is acquired, and the interactive response content is generated.
[0060] Specifically, the gesture recognition sensor is used to detect the pointing coordinates and the gesture amplitude of the zoom-in gesture.
[0061] According to the pointing coordinates, the corresponding local area center position in the target digital media object is determined, and the zoom-in multiple of the local area is determined based on the gesture amplitude. Specifically, the detected pointing coordinates such as (2.5, 1.2) are taken as the local area center position, and further, the gesture amplitude such as the double-hand zoom-in of 0.25 meters is converted into the zoom-in multiple through a preset mapping relationship. Exemplarily, the zoom-in multiple = minimum zoom-in multiple + gesture amplitude × scaling factor. If the minimum zoom-in multiple is set to 1 and the scaling factor is 2, then when the double-hand zoom-in is 0.25 meters, the zoom-in multiple = 1 + 0.25 × 2 = 1.5 times.
[0062] Based on the center position, the zoom-in multiple, and the projection size, the target local area is determined. Exemplarily, the center position is (2.5, 1.2), the zoom-in multiple is 1.5 times, and the projection size is 4 meters × 3 meters, and then the target local area can be a 0.4 meter × 0.3 meter area with (2.5, 1.2) as the center, i.e., the target local area is a rectangular area with vertex coordinates (2.3, 1.05), (2.7, 1.05), (2.7, 1.35), and (2.3, 1.35).
[0063] From the digital media file library, the third resolution data of the target digital media object is acquired. The third resolution is a higher resolution, e.g., the third resolution data of the target digital media object is 4000 × 3000.
[0064] According to the target local area, the third resolution data of the target digital media object is intercepted to obtain the third resolution data of the target local area. Specifically, the target local area of the third resolution of the target digital media image is intercepted to acquire the third resolution data of the target local area.
[0065] Based on the third resolution data of the target local area, the target local area is zoomed in and displayed, and the target local area is switched to the zoom-in display mode, e.g., the entire projection content is changed to the third resolution data of the target local area.
[0066] The zoomed-in and displayed target local area is taken as the interactive response content.
[0067] The technical effect of natural and accurate interaction is realized by detecting and recognizing diversified gesture interaction actions of a target user such as a sliding gesture, a selection gesture and a zoom-in gesture, and generating corresponding interaction response content based on real-time processing of the gestures and initial projection content. When a sliding gesture is recognized, the intention of the user to browse more can be understood, and a content set is updated to maintain continuity of interaction; when a selection gesture is recognized, a specific object pointed by the user can be accurately positioned and switched to a separate display mode to respond to the demand of the user for in-depth understanding; when a zoom-in gesture is recognized, the demand of the user for exploring details can be further met, and a specified local area is zoomed in to display, thus meeting the demand of the user for in-depth understanding. Through the resolution on-demand loading mechanism, high-resolution data is provided only when the audience really needs it, which can greatly reduce network transmission pressure and storage space demand. It is ensured that multiple people can use the device at the same time under limited hardware conditions, and the use efficiency of the device is improved.
[0068] In an embodiment, as shown in FIG. 1, the application provides an interactive processing method of a digital media file, including: Figure 2 According to the same inventive concept of the interactive processing method of the digital media file provided in the embodiment, the application further provides an interactive processing system of a digital media file, including: A user positioning acquisition module 100 is configured to detect user positioning information of a target user entering a digital interaction area, wherein the user positioning information includes a coordinate position, a visual height and a visual orientation angle of the target user in the digital interaction area; A display parameter generation module 200 is configured to generate personalized display parameters for the target user according to the coordinate position, the visual height and the visual orientation angle, wherein the personalized display parameters include a projection center point, a projection size and a projection angle; A projection content acquisition module 300 is configured to select corresponding media content in a digital media file library based on the projection center point, the projection size and the projection angle to process and obtain initial projection content, wherein the initial projection content includes a plurality of digital media objects; An interaction response module 400 is configured to respond to a gesture interaction action of the target user, perform interaction processing based on the gesture interaction action and the initial projection content, and generate interaction response content.
[0069] In an embodiment, the user positioning acquisition module 100 is further configured to: A three-dimensional coordinate system of the digital interaction area is established with a central projection device of the digital interaction area as a center point, with a preset response distance as a radius and with a preset vertical detection range as a height; Position information of the target user in the three-dimensional coordinate system is detected to obtain a coordinate position of the target user in the digital interaction area; detecting a head height position of the target user, and determining a visual height of the target user according to the head height position; detecting a head vertical orientation of the target user, and determining a visual orientation angle of the target user according to the head vertical orientation; taking the coordinate position of the target user in the digital interaction area, the visual height and the visual orientation angle as the user positioning information.
[0070] In one embodiment, the display parameter generation module 200 is further configured to: connecting the coordinate position of the target user with the position of the central projection device to determine a connecting line direction and a user viewing distance; extending from the coordinate position of the target user along the connecting line direction to the central projection device by a preset viewing distance to obtain a horizontal viewing point; determining a vertical viewing point according to the visual height, and determining a reference projection point based on the horizontal viewing point and the vertical viewing point; correcting the reference projection point by an arc line based on the visual orientation angle to obtain a projection center point; performing size lookup in a distance size mapping table according to the user viewing distance to obtain a projection size; determining a projection angle based on the visual orientation angle, wherein the projection angle and the visual orientation angle are in a complementary relationship; taking the projection center point, the projection size and the projection angle as the personalized display parameters of the target user.
[0071] In one embodiment, the projection content acquisition module 300 is further configured to: determining a preset grid layout according to the projection size, and determining a quantity requirement of digital media objects based on the preset grid layout; randomly selecting a plurality of digital media objects from the digital media object library based on the quantity requirement, the digital media objects including three-dimensional media models and planar media images; obtaining first resolution data of each digital media object as initial display data of the plurality of digital media objects; obtaining basic projection content based on the preset grid layout and the initial display data of the plurality of digital media objects; adjusting the basic projection content in position and angle based on the projection center point and the projection angle to obtain the initial projection content.
[0072] In one embodiment, the interaction response module 400 is further configured to: detecting a gesture interaction action of the target user, the gesture interaction action comprising a sliding gesture, a selecting gesture and a zooming gesture; when the sliding gesture is detected, identifying a direction of the sliding gesture, re-determining a plurality of updated digital media objects from the digital media file library according to the direction of the sliding gesture, and acquiring first resolution data of the plurality of updated digital media objects, and generating an interaction response content; when the selecting gesture is detected, identifying a digital media object pointed by the selecting gesture, determining a target digital media object, acquiring second resolution data of the target digital media object, and generating an interaction response content; wherein identifying the digital media object pointed by the selecting gesture, determining the target digital media object, acquiring the second resolution data of the target digital media object, and generating the interaction response content comprise: detecting a pointing position of the selecting gesture in the projection content; determining a corresponding grid cell according to the pointing position, and identifying a digital media object in the grid cell, and determining the identified digital media object as the target digital media object; acquiring the second resolution data of the target digital media object from the digital media file library; processing the target digital media object based on the second resolution data, and switching the target digital media object to a separate display mode; displaying the target digital media object in the separate display mode as the interaction response content; when the zooming gesture is detected, identifying a local area pointed by the zooming gesture in the target digital media object, determining a target local area, acquiring third resolution data of the target local area, and generating an interaction response content; wherein identifying the local area pointed by the zooming gesture in the target digital media object, determining the target local area, acquiring the third resolution data of the target local area, and generating the interaction response content comprise: detecting a pointing coordinate and a gesture amplitude of the zooming gesture; determining a corresponding local area center position in the target digital media object according to the pointing coordinate, and determining a zooming multiple of the local area based on the gesture amplitude; determining the target local area based on the center position, the zooming multiple and the projection size; acquiring the third resolution data of the target digital media object from the digital media file library; cutting the third resolution data of the target digital media object according to the target local area to obtain third resolution data of the target local area; zoom in and display the target local region based on third resolution data of the target local region, and switch the target local region to a zoom-in display mode; display the target local region zoomed in as the interaction response content.
[0073] In summary, the embodiments of the present application have at least the following technical effects: The present application provides an interactive processing method and system for digital media files. The method includes real-time detection of the coordinate position, visual height, and visual orientation angle of a target user in a digital interaction area, dynamic generation of personalized display parameters including a projection center point, projection size, and projection angle based on the detection, intelligent selection and processing of content from a media library based on the parameters to form an initial projection, and generation of corresponding interaction response content in response to natural gesture interaction actions of the user. The method significantly improves the personalization level, immersion of user experience, and natural fluency of interaction of the digital media interaction system. Specifically, the method establishes a three-dimensional coordinate system and accurately obtains the visual height and orientation of the user, ensuring that the generated projection center point and projection angle are optimally aligned with the visual line of the user, effectively eliminating visual distortion and content obstruction caused by improper viewing position and angle, and enabling the user to obtain the best front viewing angle customized for the user. The method dynamically adjusts the projection size based on the viewing distance of the user, ensuring that the displayed content has appropriate visual proportion and clarity at different distances, and enhancing the comfort of viewing. The method combines the personalized parameters with the selection and processing of media library content, ensuring that the initial projection content is optimized for the current user state in terms of layout and presentation, and laying a good foundation for subsequent interaction. Furthermore, the method introduces precise recognition and response to gesture interaction actions such as sliding, selecting, and zooming, enabling the user to interact with the projection content in the most intuitive and natural way, such as switching content, focusing on specific objects, or exploring details of digital media. Through an on-demand loading mechanism, high-resolution data is only provided when the user actually needs it, significantly reducing network transmission pressure and storage space requirements. At the same time, a three-level resolution progressive loading strategy ensures that multiple users can use the system simultaneously under limited hardware conditions, improving the efficiency of the device. Compared with traditional methods, the technical solution provided by the present application significantly overcomes the problems of viewing angle deviation, content distortion, and stiff interaction caused by fixed projection modes, and achieves the technical effects of dynamically sensing the user state, intelligently adapting the display content, and accurately responding to natural interaction intentions, providing a highly personalized, immersive, and smooth digital media interaction experience for users.
[0074] It should be noted that the above-mentioned embodiment sequences of the present application are merely for description only, but not for representing the advantages and disadvantages of the embodiments. And the above-mentioned embodiments of the present specification have been described. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-task processing and parallel processing are also possible or can be advantageous.
[0075] The above only describes the preferred embodiments of the present application, and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0076] The specification and drawings are merely exemplary of the present application, and any and all modifications, variations, combinations or equivalents that are within the scope of the present application should be included. Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application belong to the scope of the present application and its equivalents, the present application is intended to include these modifications and variations.
Claims
1. A method for interactive processing of a digital media file, characterized by, The method comprises: detecting user positioning information of a target user entering a digital interaction area, the user positioning information comprising coordinate position, visual height and visual orientation angle of the target user in the digital interaction area; generating personalized display parameters for the target user according to the coordinate position, the visual height and the visual orientation angle, the personalized display parameters comprising a projection center point, a projection size and a projection angle; processing corresponding media content in a digital media file library based on the projection center point, the projection size and the projection angle to obtain initial projection content, the initial projection content comprising a plurality of digital media objects; generating interactive response content by processing the gesture interaction action and the initial projection content based on the gesture interaction action of the target user in response to the gesture interaction action of the target user.
2. The method of claim 1, wherein, Detecting user positioning information of a target user entering a digital interaction area, the user positioning information comprising coordinate position, visual height and visual orientation angle of the target user in the digital interaction area, comprising: establishing a three-dimensional coordinate system of the digital interaction area with a central projection device of the digital interaction area as a center point, a preset response distance as a radius and a preset vertical detection range as a height; detecting position information of the target user in the three-dimensional coordinate system to obtain the coordinate position of the target user in the digital interaction area; detecting the head height position of the target user and determining the visual height of the target user according to the head height position; detecting the vertical orientation of the head of the target user and determining the visual orientation angle of the target user according to the vertical orientation of the head; taking the coordinate position, the visual height and the visual orientation angle of the target user in the digital interaction area as the user positioning information.
3. The method of claim 2, wherein, Generating personalized display parameters for the target user according to the coordinate position, the visual height and the visual orientation angle, the personalized display parameters comprising a projection center point, a projection size and a projection angle, comprising: connecting the coordinate position of the target user and the position of the central projection device to determine the direction of the connecting line and the user viewing distance; extending a preset viewing distance from the coordinate position of the target user along the direction of the connecting line to the central projection device to obtain a horizontal viewing point; determining a vertical viewing point according to the visual height, and determining a reference projection point based on the horizontal viewing point and the vertical viewing point; correcting the reference projection point based on the visual orientation angle to obtain a projection center point; performing size lookup in a distance size mapping table according to the user viewing distance to obtain a projection size; determining a projection angle based on the visual orientation angle, wherein the projection angle and the visual orientation angle are in a complementary relationship; taking the projection center point, the projection size and the projection angle as the personalized display parameters of the target user.
4. The method of claim 1, wherein, Processing corresponding media content in a digital media file library based on the projection center point, the projection size and the projection angle to obtain initial projection content, the initial projection content comprising a plurality of digital media objects, comprising: determine a preset grid layout according to the projection size, and determine a quantity requirement of digital media objects based on the preset grid layout; randomly select a plurality of digital media objects from the digital media file library based on the quantity requirement, the digital media objects including three-dimensional media models and planar media images; obtain first resolution data of each digital media object as initial display data of the plurality of digital media objects; obtain base projection content based on the preset grid layout and the initial display data of the plurality of digital media objects; adjust the base projection content in position and angle based on the projection center point and the projection angle to obtain the initial projection content.
5. The method of claim 1, wherein, generate interactive response content based on the gesture interaction action of the target user and the initial projection content in response to the gesture interaction action of the target user, including: detect the gesture interaction action of the target user, the gesture interaction action including a sliding gesture, a selection gesture, and a zoom-in gesture; when the sliding gesture is detected, identify a direction of the sliding gesture, re-determine a plurality of updated digital media objects from the digital media file library according to the direction of the sliding gesture, obtain first resolution data of the plurality of updated digital media objects, and generate interactive response content; when the selection gesture is detected, identify a digital media object pointed by the selection gesture, determine a target digital media object, obtain second resolution data of the target digital media object, and generate interactive response content; when the zoom-in gesture is detected, identify a local area pointed by the zoom-in gesture in the target digital media object, determine a target local area, obtain third resolution data of the target local area, and generate interactive response content.
6. The method of claim 5, wherein, identify a digital media object pointed by the selection gesture, determine a target digital media object, obtain second resolution data of the target digital media object, and generate interactive response content, including: detect a pointing position of the selection gesture in the projection content; determine a corresponding grid cell according to the pointing position, and identify a digital media object in the grid cell, and determine the identified digital media object as the target digital media object; obtain second resolution data of the target digital media object from the digital media file library; process the target digital media object based on the second resolution data, and switch the target digital media object to a separate display mode; display the target digital media object in the separate display mode as the interactive response content.
7. The method of claim 5, wherein, identify a local area pointed by the zoom-in gesture in the target digital media object, determine a target local area, obtain third resolution data of the target local area, and generate interactive response content, including: detect a pointing coordinate and a gesture amplitude of the zoom-in gesture; determine a corresponding local area center position in the target digital media object according to the pointing coordinate, and determine a zoom-in multiple of the local area based on the gesture amplitude; determine the target local area based on the center position, the zoom-in multiple, and the projection size; obtain third resolution data of the target digital media object from the digital media file library; According to the target local area, the third resolution data of the target digital media object is intercepted to obtain third resolution data of the target local area; Based on the third resolution data of the target local area, the target local area is zoomed in and displayed, and the target local area is switched to zoomed-in display mode; The zoomed-in target local area is used as the interactive response content.
8. An interactive processing system for digital media files, characterized by The system for implementing the method of any one of claims 1-7 comprises: A user position acquisition module for detecting user position information of a target user entering a digital interaction area, the user position information including the coordinate position, visual height and visual orientation angle of the target user in the digital interaction area; A display parameter generation module for generating personalized display parameters for the target user according to the coordinate position, visual height and visual orientation angle, the personalized display parameters including a projection center point, projection size and projection angle; A projection content acquisition module for selecting corresponding media content from a digital media file library based on the projection center point, projection size and projection angle to obtain initial projection content, the initial projection content including a plurality of digital media objects; An interactive response module for responding to gesture interaction actions of the target user, performing interactive processing based on the gesture interaction actions and the initial projection content to generate interactive response content.
Citation Information
Patent Citations
Projection parameter determining method and device based on indoor positioning and projection system
CN112416135A
Non-contact interaction method and device, electronic equipment and medium
CN112527110A
Gesture control method and device based on display screen, equipment and storage medium
CN118409661A
Historic building protection display method and system based on virtual reality
CN119556822A
Region-limited virtual touch interaction method and related equipment
CN120295487A