A large-screen positioning method and system based on a hidden code and a smart phone

CN122526437APending Publication Date: 2026-08-07SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANDONG UNIV
Filing Date
2026-07-08
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

智能手机作为观众随身携带的通用设备,是大屏交互的理想终端,但现有手机交互方案存在明显不足:部分方案需要观众低头查看手机屏幕完成操作,无法持续注视大屏,破坏观看连续性;部分方案依赖手机传感器(陀螺仪、磁力计、加速度计)实现指向,易受环境磁场、手机姿态干扰,定位误差大;部分方案采用屏幕扫码识别可见标签,标签会遮挡大屏显示内容,影响视觉美观度与内容呈现效果;部分方案需要手机与大屏进行高精度时序同步,受手机操作系统非实时性、无线传输延迟影响,难以实现工程化落地

Benefits of technology

本发明利用人眼视觉系统的闪烁融合效应,在高刷新率屏幕上通过帧交替显示与RGB通道循环移位策略,将定位标识以隐码形式嵌入显示内容,实现标识对人眼完全不可见;同时,通过特定的色调分离与偏移增强算法,使智能手机摄像头能够从帧混叠图像中稳定提取隐码,进而完成大屏指向定位。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122526437A_ABST
    Figure CN122526437A_ABST
Patent Text Reader

Abstract

The application belongs to the field of human-computer interaction, and provides a large-screen positioning method and system based on a hidden code and a smart phone. A hidden code identifier is embedded in original display content of a large screen. In response to a shooting instruction of the smart phone, an image containing a display picture of the target large screen is acquired. The embedded hidden code identifier previously hidden in the image is extracted and identified. The projection position of the camera optical center of the smart phone on the target large screen plane is calculated. A cursor icon is displayed on the projection position. The pointing position of the smart phone on the large screen is determined through the moving track of the cursor icon. In response to the operation of the smart phone pointing at the large screen, the cursor icon is controlled to follow the movement in real time according to the determined pointing position. In response to the selection or control instruction of the user based on the cursor icon, corresponding selection or control operation is performed on the large screen. The application supports multi-user concurrent operation at a long distance, and realizes accurate large-screen positioning interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of human-computer interaction, specifically relating to a method and system for large-screen positioning based on hidden codes and smartphones. Background Technology

[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.

[0003] With the rapid development of LED display technology, projection blending technology, and extended reality technology, cave-like XR display environments, circular public screens, and immersive exhibition screens composed of single or multiple ultra-large screens have been widely used in museums, science and technology museums, schools, medical and health care facilities, and commercial exhibitions. These large-screen display systems can provide audiences with immersive and visually impactful multimedia information presentation effects, supporting multiple viewers to watch and learn together in the same physical space. They are important hardware carriers for future smart displays, immersive education, and digital culture dissemination.

[0004] However, current large public displays and cave-style XR display environments have significant technical shortcomings in terms of human-computer interaction, specifically in the following aspects: Close-range touch interaction fails. Ultra-large screens require viewers to maintain a viewing distance of several meters or even more than ten meters to fully grasp the overall content displayed. Traditional close-range contact interaction methods such as capacitive touch, infrared touch, and resistive touch cannot be used in long-distance viewing scenarios. Viewers cannot directly select, locate, zoom, or perform other operations on the screen content. Large screens can only achieve one-way information playback and cannot meet personalized interaction needs.

[0005] Dedicated interactive devices are costly to deploy and have poor applicability. Existing large-screen remote interactive technologies mostly rely on additional hardware devices, such as motion capture systems, eye-tracking devices, LiDAR, depth cameras, infrared positioning modules, and dedicated handles. These devices are not only expensive to procure and deploy, and complex to install and wire, but also impose strict limitations on exhibition spaces, viewing routes, and ambient lighting, making them difficult to adapt to open, large-space, high-traffic application scenarios such as museums and science museums. Furthermore, dedicated devices require viewers to wear or hold additional devices, increasing their burden and reducing the naturalness and convenience of the interactive experience.

[0006] The stability and accuracy of aerial gesture interaction are insufficient. Vision-based aerial gesture interaction, which requires no physical contact with the device, is a hot research topic for large-screen remote interaction. However, this technology is susceptible to interference from ambient lighting, audience limbs obstructing the view, and multiple simultaneous operations. Gesture recognition suffers from high latency, high false touch rate, and low positioning accuracy, making it impossible to accurately select and locate screen content. Furthermore, gesture interaction requires viewers to learn specific action commands, increasing cognitive load and making it unsuitable for viewers of all ages and without technical backgrounds.

[0007] Traditional smartphone-based interaction methods suffer from user experience flaws. While smartphones, as ubiquitous devices carried by viewers, are ideal terminals for large-screen interaction, existing smartphone interaction solutions have significant shortcomings: some solutions require viewers to look down at their phone screens to complete operations, disrupting the viewing continuity as they cannot maintain continuous eye contact with the large screen; some solutions rely on phone sensors (gyroscopes, magnetometers, accelerometers) for pointing, which are easily affected by environmental magnetic fields and phone posture, resulting in large positioning errors; some solutions use screen scanning to identify visible labels, which can obscure the content displayed on the large screen, affecting visual aesthetics and content presentation; and some solutions require high-precision time synchronization between the phone and the large screen, which is difficult to implement in practice due to the non-real-time nature of the phone's operating system and wireless transmission latency.

[0008] The system lacks support for concurrent multi-user interaction. Most existing large-screen interactive systems operate in a single-user control mode, where instructors and administrators switch content uniformly via consoles or tablets. This fails to support multiple viewers simultaneously in the same space to independently locate, select, and query content on the large screen according to their own needs, making it difficult to meet the core requirements of multi-person collaborative learning and personalized browsing.

[0009] To address the aforementioned technical challenges, the industry has attempted to implement large-screen interaction using visual marker-assisted positioning. This involves displaying visual identifiers such as QR codes, April Tags, and SpotCodes on the large screen, which are then captured and identified using a mobile phone camera for positioning. However, these identifiers are all visible, which can cover and interfere with the original content displayed on the large screen, severely impacting the viewer's visual experience. Some hidden marker solutions require viewers to wear auxiliary hardware such as LCD shutter glasses and special filters, increasing usage costs and operational complexity, and failing to achieve seamless interaction. Summary of the Invention

[0010] To address the aforementioned problems, this invention proposes a large-screen positioning method and system based on hidden codes and smartphones. This invention requires no additional hardware, does not affect the display effect, supports multi-user long-distance concurrent operation, and achieves accurate large-screen positioning and interaction.

[0011] According to some embodiments, the present invention adopts the following technical solution: A method for locating large-screen smartphones based on coded information includes the following steps: Embed hidden code identifiers into the original display content on the large screen, generate hidden code identifier frames, difference compensation frames, and RGB channel cyclic shift frames, and construct a cyclic frame sequence that meets the requirements of flicker fusion, so that the hidden code identifiers are invisible to the human eye and only the original display content is displayed; In response to the smartphone's shooting command, an image containing the display screen of the target large screen is acquired, and embedded hidden code identifiers are extracted and identified in the image. Based on the transformation relationship between the smartphone's camera coordinate system and the world coordinate system, the projection position of the smartphone's camera optical center on the target large screen plane is calculated, and a cursor icon is displayed at the projection position. The pointing position of the smartphone on the large screen is determined by the movement trajectory of the cursor icon. In response to the smartphone pointing at the large screen, the cursor icon moves in real time according to the determined pointing position, and the corresponding selection or control operation is performed on the large screen in response to the user's selection or control command based on the cursor icon.

[0012] As an alternative implementation, hidden code identifiers are embedded in several low-texture, color-flat areas of the original display image. Three hidden code embedding frames, F1, F2, and F3, are generated based on RGB channel cyclic shifting, and an original image compensation frame, F4, is generated to form a four-frame cyclic sequence at a predetermined frequency. This cyclic sequence is used to control the alternating playback sequence of the hidden code frame, compensation frame, and original frame according to a set refresh rate, ensuring that the hidden code is invisible to the human eye and that the human eye only sees a continuous and complete original display image.

[0013] As a further defined implementation, the process of generating three hidden code embedding frames F1, F2, and F3 based on RGB channel cyclic shifting, and generating the original image compensation frame F4, includes: using the B channel of the original displayed image as the screen background; displaying the hidden code white pixels using the R channel; and displaying the hidden code black pixels using the G channel, thereby generating a blue background hidden code frame. frame; The R channel of the original displayed image is used as the background; the hidden white pixels are displayed using the G channel; and the hidden black pixels are displayed using the B channel, generating a red background hidden frame. frame; The G channel of the original displayed image is used as the background; the hidden white pixels are displayed using the B channel; and the hidden black pixels are displayed using the R channel, generating a green background hidden frame. frame; It directly displays the original color image or original grayscale image, without embedding any hidden identifiers, and is only used for image brightness compensation, color restoration, and visual comfort enhancement, generating original image compensation frames. frame.

[0014] As an alternative implementation, the hidden identifier is in , , In the three frames, the hidden code is embedded into the image through an RGB channel cyclic shift strategy, creating a significant tonal difference between the hidden code and the background. The frames do not carry hidden code information and are used to offset the color shift and brightness decay of the first three frames. The four frames are output continuously in a fixed order. The human eye cannot perceive the content of a single frame and can only observe the original image after fusion without any markings, thus achieving complete visual hiding of the hidden code.

[0015] As an alternative implementation, in response to a smartphone's shooting command, an image containing the display screen of the target large screen is acquired. The process of extracting and identifying the pre-hidden embedded steganographic identifier in the image includes: converting the steganographic identifier image captured by the smartphone from the RGB color space to the HSV color space, extracting the hue H channel, normalizing the value of the hue H channel, generating three hue offset copies, converting the three copies into grayscale images and performing binarization and morphological denoising, simultaneously performing steganographic identification on the three hue offset copies, counting the number of successful identifications for each copy, selecting the copy with the most identifications as the final valid steganographic data, identifying and decoding each complete steganographic identifier structure contained in the image, obtaining the unique ID code information corresponding to each steganographic identifier structure and its spatial pose transformation data relative to the camera coordinate system.

[0016] As a further defined implementation, the process of extracting the hue H channel, normalizing the value of the hue H channel, and generating three hue offset copies includes: normalizing the value of the hue H channel to the [0, 1] interval to eliminate the numerical ambiguity of the red tone at 0° and 360°. For the normalized H-channel image, three independent tone-shifted copies are generated, where the H-channel of the first copy remains unchanged, denoted as . The second copy's H channel is offset by a certain value M, denoted as... ,like Then update The H channel of the third copy is offset by 2M, denoted as... ,like Then update .

[0017] As a further defined implementation method, the process of counting the number of successful identifications for each copy and selecting the copy result with the most identifications as the final valid hidden code data includes: performing hidden code detection and decoding in parallel on the binarized candidate image of each copy, that is, performing corner detection, edge extraction and ID decoding on each copy; counting the number of hidden codes successfully identified for each copy; selecting the copy result with the most identifications as the final valid hidden code data, and if multiple copies have the same number of identifications, selecting the result with the highest confidence.

[0018] As an alternative implementation, the process of calculating the projection position of the smartphone camera's optical center on the target large screen plane based on the transformation relationship between the smartphone's camera coordinate system and the world coordinate system includes: establishing a world coordinate system with the endpoint of the large screen as the origin, where the X-axis of the world coordinate system is the horizontal direction of the screen, the Y-axis is the vertical direction of the screen, and the Z-axis is perpendicular to the screen plane and points outward; calculating a set of candidate projection points of the smartphone's optical center on the screen plane based on the spatial pose transformation data of the hidden identifier relative to the camera coordinate system and the transformation relationship between the smartphone's camera coordinate system and the world coordinate system; then performing interpolation calculations on the candidate projection points to finally obtain the precise coordinates of the smartphone's optical center mapped on the screen.

[0019] As an alternative implementation, the hidden code identifier is a QR code.

[0020] A large-screen positioning system based on coded information and smartphones includes: The display-end hidden code embedding module is used to embed hidden code identifiers into the original display content of the large screen, generate hidden code identifier frames, difference compensation frames, and RGB channel cyclic shift frames, and construct a cyclic frame sequence that meets the requirements of flicker fusion, so that the hidden code identifiers are invisible to the human eye and only the original display content is displayed. The mobile phone pointing and positioning calculation module is used to respond to the shooting command of the smartphone, acquire an image containing the display screen of the target large screen, extract and identify the embedded hidden code identifiers in the image, calculate the projection position of the optical center of the smartphone camera on the plane of the target large screen according to the transformation relationship between the smartphone's camera coordinate system and the world coordinate system, display a cursor icon at the projection position, and determine the pointing position of the smartphone on the large screen by the movement trajectory of the cursor icon. The human-computer interaction module is used to respond to the operation of the smartphone pointing at the large screen. According to the determined pointing position, it controls the cursor icon to move in real time and responds to the user's selection or control command based on the cursor icon to perform the corresponding selection or control operation on the large screen.

[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention utilizes the flicker fusion effect of the human visual system to embed the positioning mark into the display content in the form of a hidden code on a high refresh rate screen through frame alternation display and RGB channel cyclic shifting strategy, making the mark completely invisible to the human eye; at the same time, through specific tone separation and offset enhancement algorithms, the smartphone camera can stably extract the hidden code from the frame aliasing image, thereby completing the large screen pointing positioning.

[0022] This invention achieves a truly imperceptible hiding effect without disrupting the display by utilizing the human eye flicker fusion effect: based on the human eye flicker fusion effect and the alternating display of RGB channel cyclic shift frames, the hidden code is completely invisible to the human eye, without obstructing, interfering with, or changing the original display content, thus ensuring the best viewing experience; This invention achieves stable extraction of hidden codes by mobile phones through a color space conversion method: by using a hue three-copy offset enhancement algorithm, it overcomes the frame aliasing interference generated by a mobile phone shooting a 240Hz screen at 120 frames per second (FPS), and can stably extract hidden codes regardless of the phase of the aliasing. The solution of this invention requires no additional hardware and has extremely low deployment costs: it does not require any auxiliary hardware such as motion capture, radar, infrared sensors, LCD shutter glasses, or special filters, and can be achieved solely with a 240Hz high refresh rate screen and a smartphone that supports 120Hz shooting. The solution of this invention does not require dual-end timing synchronization and has low engineering difficulty: it does not require frame synchronization signal transmission between the large screen and the mobile phone, and is not affected by the non-real-time operating system of the mobile phone, wireless latency, or system scheduling jitter, making it simple and reliable to implement. The method of this invention is a content-independent pointing and positioning method, which effectively improves the versatility and practicality of the technology. In this invention, the QR code used for positioning is an independent display frame, completely unrelated to the original content displayed on the screen. It eliminates the need for prior feature extraction and registration of the original image, and also eliminates the need to download the feature information of the displayed image to the mobile device in advance. This simplifies the interaction process, breaks through the dependence of traditional positioning methods on the content of the displayed image, and expands the applicability of the pointing and positioning function.

[0023] This invention is based on large-screen exhibitions and presents diverse content in various formats, including videos and audio narration. The diversified content presentation attracts user interest and enhances user engagement and immersion. Simultaneously, the exhibition content is highly flexible and adaptable to different application scenarios. It is not only suitable for museum and art gallery exhibitions but can also be widely applied to various venues such as school education, community cultural activities, and commercial displays, thus significantly expanding the system's application scope.

[0024] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0025] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.

[0026] Figure 1 Timing diagram for embedding and displaying hidden codes in a four-frame loop for a 240Hz screen; Figure 2 The diagram shows four sampling scenarios of 240Hz screen frame aliasing when shooting with a 120FPS mobile phone. (a) represents sampling scenario one and sampling scenario three, and (b) represents sampling scenario two and sampling scenario four. Figure 3 Four aliasing sampling modes and linear interpolation models are used; Figure 4 This is a schematic diagram of a method for extracting hidden codes for aliased image tonal three-copy offset enhancement according to one embodiment; Figure 5 This is a schematic diagram of aliased image tone three-copy offset enhancement hidden code extraction according to another embodiment. Detailed Implementation

[0027] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0028] It should be noted that the following detailed description is illustrative and intended to provide further explanation of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0029] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0030] Where there is no conflict, the embodiments and features described in this application may be combined with each other.

[0031] Example 1 This embodiment provides a large-screen positioning method based on hidden codes and smartphones. Its core principle is that the human eye's flicker fusion frequency threshold is 50Hz–60Hz. When the equivalent refresh rate of intermittent flickering light reaches this threshold, the human eye perceives it as a continuous and stable light signal. Based on this principle, on a high refresh rate screen such as 240Hz, the original display image is decomposed into multiple alternating frames. The first few frames embed hidden code identifiers using a cyclic shifting method of RGB channels, while the last frame is the original image used for brightness compensation. In the human visual system, these multiple frames are rapidly and alternately fused into a complete and continuous original image, completely hiding the hidden code. Simultaneously, when a smartphone captures the screen at 120FPS, although frame aliasing occurs, the hidden code identifier can be stably separated by extracting the tone channels of the aliased image, enhancing the tone shift in three directions, and selecting the optimal result, thus achieving reliable extraction and recognition of the hidden code.

[0032] A method for locating large-screen smartphones based on coded information includes the following steps: Embed hidden code identifiers into the original display content on the large screen, generate hidden code identifier frames, difference compensation frames, and RGB channel cyclic shift frames, and construct a cyclic frame sequence that meets the requirements of flicker fusion, so that the hidden code identifiers are invisible to the human eye and only the original display content is displayed; In response to the smartphone's shooting command, an image containing the display screen of the target large screen is acquired, and embedded hidden code identifiers are extracted and identified in the image. Based on the transformation relationship between the smartphone's camera coordinate system and the world coordinate system, the projection position of the smartphone's camera optical center on the target large screen plane is calculated, and a cursor icon is displayed at the projection position. The pointing position of the smartphone on the large screen is determined by the movement trajectory of the cursor icon. In response to the smartphone pointing at the large screen, the cursor icon moves in real time according to the determined pointing position, and the corresponding selection or control operation is performed on the large screen in response to the user's selection or control command based on the cursor icon.

[0033] Specifically, in this embodiment, a hidden code identifier is embedded in the low-texture, color-flattened area of ​​the original displayed image. Three hidden code embedding frames (F1, F2, and F3) are generated based on RGB channel cyclic shifting, and an original image compensation frame (F4) is generated, forming a 240Hz four-frame cyclic sequence. This cyclic sequence is used to control the alternating playback order of the hidden code frame, compensation frame, and original frame at high refresh rates such as 120Hz and 240Hz, ensuring that the hidden code is invisible to the human eye. The image is then played in a 240Hz refresh rate loop. , , , The frame sequence utilizes the human eye flicker fusion effect to make the hidden code invisible to the human eye, so that the human eye only sees the continuous and complete original display image.

[0034] In this embodiment, the hidden code identifier can be a QR code planar marking pattern, such as AprilTag, ArUco, ARTag, etc.

[0035] When a mobile phone captures this type of QR code, the corresponding open-source code can be used to identify the QR code's ID number and calculate the transformation relationship between the phone's camera coordinate system and the world coordinate system.

[0036] In the large-space public display environment used in this embodiment, the display device uses a world coordinate system. For example, the lower left corner of the screen is taken as the origin of the world coordinate system. The X-axis of the world coordinate system is the horizontal direction of the screen, the Y-axis is the vertical direction of the screen, and the Z-axis is perpendicular to the screen plane and points outward. At this time, the projection position of the optical center of the mobile phone camera on the screen plane can be calculated based on the QR code image captured by the mobile phone in real time. A cursor icon is then displayed at this projection position. When the user holds the mobile phone to scan and capture the screen image, the movement trajectory of the cursor icon displayed on the screen can be used to determine the pointing position of the mobile phone on the screen, which serves as the basis for interactive operations such as target selection.

[0037] In terms of specific implementation, the QR code image captured by the mobile phone is first subjected to adaptive binarization processing to effectively remove the interference of stray light and noise in the image on the QR code structure. Then, through the corresponding open source code, each complete QR code structure contained in the image is identified and decoded to obtain the unique ID code information corresponding to each QR code structure, as well as its spatial pose transformation data relative to the mobile phone camera coordinate system. Based on this, a set of candidate projection points of the mobile phone optical center on the screen plane is calculated. Then, interpolation calculation is performed on the candidate projection points to finally obtain the accurate coordinates of the mobile phone optical center mapped on the screen.

[0038] In response to a user's gesture of pointing their phone at a large public display screen, the system controls the interactive icons on the screen to move in real time according to the final mapped coordinates. At the same time, it supports users in selecting and controlling videos, images, and other functions on the screen based on these coordinates, and allows users to create interactive effects such as drawing and recording traces on the screen via their mobile phones.

[0039] Example 2 A large-screen positioning system based on coded information and smartphones includes: a display-end coded information embedding module, a mobile-end positioning interaction processing module, and a human-computer interaction module; The display-end hidden code embedding module is used to embed hidden code identifiers into the original display content of the large screen, generate hidden code identifier frames, difference compensation frames, and RGB channel cyclic shift frames, and construct a frame sequence that meets the flicker fusion requirements; it embeds hidden code identifiers in low-texture, color-flat areas of the original display image, and generates them based on RGB channel cyclic shift. , , Three hidden code embedding frames, and generation The original image compensation frames form a 240Hz four-frame loop sequence; this is used to control the alternating playback sequence of the hidden code frames, compensation frames, and original frames according to high refresh rates such as 120Hz and 240Hz, ensuring that the hidden code is invisible to the human eye; it loops at a 240Hz refresh rate. , , , The frame sequence utilizes the human eye flicker fusion effect to make the hidden code invisible to the human eye, so that the human eye only sees the continuous and complete original display image; The aforementioned mobile phone pointing and positioning calculation module essentially involves a smartphone capturing images of the screen at 120 FPS, obtaining images of frame aliasing caused by frame rate mismatch; converting the captured images from RGB color space to HSV color space, extracting the hue H channel; generating three hue offset copies: H unchanged, H+0.333, and H+0.666, and normalizing hue values ​​outside the range by modulo operation; converting the three copies into grayscale images and performing binarization and morphological denoising; simultaneously performing hidden code recognition on the three hue offset copies, counting the number of successful recognitions for each copy, and selecting the copy with the most successful recognitions as the final valid hidden code data; subsequently, using the corresponding open-source code, recognizing and decoding each complete QR code structure contained in the image, obtaining the unique ID code information corresponding to each QR code structure, as well as its spatial pose transformation data relative to the mobile phone camera coordinate system, and calculating a set of candidate projection points of the mobile phone's optical center on the screen plane, and then performing interpolation calculations on the candidate projection points to finally obtain the precise coordinates of the mobile phone's optical center mapped on the screen.

[0040] The human-computer interaction module is used to respond to the user's operation of pointing their mobile phone at a large public display screen. Based on the final mapped coordinates output by the positioning calculation module on the mobile phone, it controls the interactive icons on the large screen to move in real time. At the same time, it supports the user to select and control videos, pictures and other functions on the large screen based on the coordinates, and can also realize interactive effects such as drawing and trajectory recording on the large screen through mobile phone operation.

[0041] In the specific implementation process, the display terminal hidden code embedding module is used to: plan and lay out the hidden code markings in the display screen according to the screen display area size, resolution, and viewing distance parameters; and decompose the original display screen into four frames according to the 240Hz high refresh rate screen four-frame timing display rules. , , , The hidden code is embedded into the corresponding frame image by cyclically shifting the RGB channels; , , , The fixed-time loop output display screen utilizes the human eye flicker fusion effect to make the hidden code invisible to the human eye, while ensuring that the hidden code can be recognized by the mobile phone camera, providing a stable identification foundation for subsequent mobile phone shooting and recognition.

[0042] In this embodiment, the mobile phone positioning and interaction processing module is used to: capture the large screen display at a frame rate of 120FPS to obtain an RGB image containing frame aliasing; convert the captured image to the HSV color space and extract the hue channels, and complete the normalization process; generate three sets of hue offset copies: H invariant, H+0.333, and H+0.666; convert the three sets of hue images to grayscale images and perform binarization and morphological denoising processing; perform hidden code recognition and decoding on the three sets of images in parallel, select the result with the best recognition effect, and output the hidden code ID and corner coordinates; and, based on the pinhole imaging and perspective projection model, complete the mobile phone pose calculation and large screen pointing coordinate calculation, and output accurate interactive positioning coordinates.

[0043] In this embodiment, the human-computer interaction module is used to: receive the large-screen pointing coordinate data output by the mobile phone's positioning interaction processing module, and map the coordinates to the corresponding interactive area of ​​the large-screen display; perform operations such as selecting, confirming, and viewing details of the displayed content based on the pointing coordinates, supporting interaction forms such as single-finger pointing selection, long-press confirmation, and swiping switching; synchronously feed back user interaction commands to the display content control system, update the screen display content in real time, and realize real-time interaction between users and virtual exhibits and display information on the large screen; support multiple users to initiate interactions independently at the same time without interference or crosstalk, meeting the parallel interaction needs in scenarios such as multi-person learning, immersive exhibitions, and interactive experiences.

[0044] In this embodiment, the basic equipment includes a large display screen that supports a 240Hz refresh rate and a smartphone that supports 120Hz shooting.

[0045] Figure 1 This embodiment illustrates the timing diagram of the four-frame cyclical hidden code embedding and display on a 240Hz screen. On a 240Hz high refresh rate display, this invention divides the displayed content into a cyclical display sequence of four frames, with each frame lasting 1 / 240 of a second. The total display time for the four frames is approximately 16.67 milliseconds, resulting in an equivalent perceived refresh rate of 60Hz.

[0046] The process includes the following steps: Step S101: 240Hz screen four-frame loop steganography embedding rules Step S1011: Frame: Blue background coded frame The blue channel (B channel) of the original displayed image is used as the background; the hidden white pixels are displayed using the red channel (R); and the hidden black pixels are displayed using the green channel (G).

[0047] Step S1012: Frame: Red background hidden frame The red channel (R channel) of the original displayed image is used as the background; the hidden white pixels are displayed using the green channel (G); and the hidden black pixels are displayed using the blue channel (B).

[0048] Step S1013: Frame: Green background hidden frame The green channel (G channel) of the original display image is used as the background; the hidden white pixels are displayed using the blue channel (B); and the hidden black pixels are displayed using the red channel (R).

[0049] Step S1014: Frame: Original image compensation frame It directly displays the original color image or the original grayscale image, without embedding any hidden codes, and is only used for screen brightness compensation, color restoration and visual comfort improvement.

[0050] exist , , In all three frames, the hidden code is embedded into the image through a cyclic shift strategy of the RGB channel, so that there is a significant color difference between the hidden code and the background, which makes it easier for the mobile phone camera to separate and extract it. The first frame does not carry hidden code information and is used to offset the color shift and brightness decay of the first three frames. The four frames are output continuously in a fixed order. The human eye cannot perceive the content of a single frame and can only observe the complete, natural, and unmarked original image after fusion, thus achieving complete visual concealment of the hidden code.

[0051] Step S102: Stealth principle based on scintillation fusion effect The human visual system has a critical flicker fusion frequency (CFF), which is approximately 50Hz–60Hz for the average adult. When light or images are displayed intermittently at frequencies higher than this, the human eye cannot distinguish the flicker and will automatically fuse them into a continuous, stable, flicker-free image.

[0052] This embodiment utilizes this physiological characteristic to achieve code hiding: the 240Hz screen has a display cycle of four frames, and the equivalent perceived refresh rate is: 240Hz ÷ 4 = 60Hz, which falls exactly within the human eye's flicker fusion range.

[0053] Under these conditions: , , The embedded code switches rapidly with the high-speed frame sequence and is averaged in brightness, canceled in color, and fused in outline in the human visual system. The human eye cannot perceive the hidden code structure in a single frame and can only see the original display content formed by the fusion of four frames. Ultimately, the technical effect of making the hidden code completely invisible to the human eye, without destroying the image, and without affecting the viewing experience is achieved.

[0054] Step S103: Frame timing drive control This embodiment uses a purely asynchronous driving mechanism, which does not require any form of timing synchronization with the smartphone.

[0055] The frame timing driver module outputs image frames at a fixed frequency of 240Hz, and the output order is strictly as follows: , , , , , , , ... The system operates entirely asynchronously, requiring no frame synchronization pulses, no wireless broadcast timing signals, no hardware triggering, and no time alignment between the mobile phone and the large screen. This design significantly reduces system complexity, engineering deployment difficulty, and hardware costs, while improving stability and versatility.

[0056] Figure 2 (a) and Figure 2 of (b) Figure 3 The illustrations show four sampling scenarios of 240Hz screen frame aliasing captured by a 120FPS mobile phone. Since the maximum shooting rate of mainstream smartphone cameras is 120FPS, which is lower than the 240Hz display refresh rate of large screens, the screen will display two complete frames plus part of a third frame within the exposure time of one frame of the phone image, resulting in frame aliasing.

[0057] The aliased captured image A can be uniformly represented as a linear weighted combination of three consecutive screen frames: ; Where t∈[0,1] are linear interpolation coefficients, which are determined by the difference between the phase at which the phone starts shooting and the phase displayed on the screen; after shooting begins, t remains constant for a short period of time.

[0058] Based on the four-frame loop sequence of the screen, there are only four sampling scenarios in mobile phone shooting: Sampling Scenario 1: ; Sampling Scenario 2: ; Sampling Scenario 3: ; Sampling Scenario 4: ; The core breakthrough of this invention lies in the fact that, regardless of the value of t or the current sampling scenario, the hidden code structure can be stably and accurately extracted from the aliased image through the tone three-copy offset enhancement algorithm, thereby achieving reliable recognition without hardware synchronization, shutter device, or timing calibration.

[0059] Figure 4 , Figure 5 This invention demonstrates a schematic diagram of enhanced occult code extraction from aliased images using a smartphone with tonal three-copy offset. The smartphone client of this invention only performs the necessary operations required for occult code extraction and does not include non-core logic such as cursor rendering, interface control, multi-device networking, and extended interaction. This achieves reliable occult code extraction from aliased images with lightweight design, high real-time performance, and high robustness.

[0060] The process includes the following steps: Step S201: RGB to HSV color space conversion The acquired aliased RGB image is converted to the HSV color space, and only the hue (H) channel is retained, while the saturation (S) and luminance (V) channels are discarded. This eliminates the luminance interference caused by ambient light, screen brightness, and shooting angle, and retains only the most stable hue difference information.

[0061] The hue H is calculated based on the segments where the maximum values ​​of R, G, and B are located, ensuring a continuous and uniform hue angle distribution.

[0062] The core advantage of using the HSV color space is that the tonal difference between the hidden code and the background is light-invariant, unaffected by ambient light, screen brightness, shooting distance and angle, and can stably distinguish between hidden code and background pixels in complex scenes.

[0063] Step S202: Tone Channel Normalization The values ​​of the hue H channel are normalized to the range of [0, 1], eliminating the ambiguity of red hue values ​​at 0° and 360°, making the hue distribution continuous and uniform, and directly usable for subsequent grayscale mapping. The normalized hue values ​​can directly reflect the pixel color category and are not affected by the numerical jumps caused by angle cycles.

[0064] Step S203: Generate three tone-off copies For the normalized H-channel image, three independent tone offset copies are generated to cover four frame aliasing scenarios, ensuring that the hidden code can be correctly extracted regardless of the aliasing phase.

[0065] Three replica generation rules: Copy 1: The H channel remains unchanged, denoted as ; Copy 2: Overall offset of H channel +0.333 (corresponding to 120°), denoted as ;like ,but ; Copy 3: Overall offset of H channel +0.666 (corresponding to 240°), denoted as ;like ,but .

[0066] The core function of this step is that, regardless of which of the sampling scenarios 1–4 the current frame aliasing is in, one of the three copies will have a tone distribution that perfectly matches the original black and white structure of the hidden code, and can be decoded normally by the standard recognizer, thus ensuring the success rate of extraction in principle.

[0067] Step S204: Convert the image from hue to grayscale. Convert the three tone-off copies to grayscale images respectively, according to the following rules: The hue value is directly linearly mapped to the grayscale value, where 0 corresponds to pure black and 1 corresponds to pure white. The black and white structure of the hidden code is completely restored without brightness or saturation interference. The converted grayscale image can clearly present the outline, corners and coding structure of the hidden code, which is consistent with the standard QR code image and can be directly sent to a general decoder for recognition.

[0068] Step S205: Binarization and Morphological Denoising The following processing steps are performed sequentially on the grayscale image to obtain a clean binarized steganographic image: fixed threshold binarization to strictly separate the steganographic code from the background; morphological opening operation to remove isolated noise points and small spikes; and connected component area filtering to remove excessively small regions, retaining only the complete structure of the steganographic code. After processing, three copies are output as a high-quality binarized steganographic code candidate image.

[0069] Step S206: Parallel Identification and Optimal Selection of Three Replicas Hidden code detection and decoding are performed in parallel on three binarized candidate images: corner detection, edge extraction, and ID decoding are performed on each copy; the number of hidden codes successfully identified in each copy is counted; the copy with the most identified hidden codes is selected as the final valid hidden code data; if multiple copies have the same number of identified hidden codes, the result with the highest confidence is selected.

[0070] At this point, the smartphone completes the entire hidden code extraction process and outputs the processed image content. Next, it performs screen pointing location calculations. The pointing location principle of this invention essentially involves reasonably transplanting and adapting the mature "planar marker pose estimation" technology from computer vision to the pointing location application scenario of "large screen + mobile phone." By visually embedding and hiding the QR code image, it becomes invisible to the human eye, ensuring that the QR code image, as a planar marker, has no impact on the viewing effect, and the mobile phone can capture the hidden QR code image.

[0071] In summary, to enable users to obtain a more intuitive, richer, and more stable immersive interactive experience in public displays, science exhibitions, immersive learning, and multi-person interactive scenarios, this embodiment adopts a large-screen positioning interaction method based on hidden codes, which effectively solves the technical problems of traditional large-screen display systems having a single interaction form, only supporting one-way information presentation, and being unable to support multiple people's simultaneous personalized operation.

[0072] Unlike existing interaction methods that rely solely on visible markers, dedicated hardware, or smartphone touch swipes, this embodiment proposes a virtual-real fusion positioning method based on invisible hidden codes and smartphone visual recognition. This system allows users to simultaneously view content and select objects on the same display screen, ensuring spatial consistency between interactive input and display output. Furthermore, users only need their smartphones to achieve long-distance, real-time, and precise interaction with large-screen content, eliminating the need for additional equipment such as shutter glasses or data gloves, or dedicated hardware like infrared, radar, or motion capture systems. This significantly improves the convenience, practicality, and applicability of the interaction system.

[0073] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made by those skilled in the art without creative effort within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for positioning large-screen smartphones based on coded information, characterized in that, Includes the following steps: Embed hidden code identifiers into the original display content on the large screen, generate hidden code identifier frames, difference compensation frames, and RGB channel cyclic shift frames, and construct a cyclic frame sequence that meets the requirements of flicker fusion, so that the hidden code identifiers are invisible to the human eye and only the original display content is displayed; In response to the smartphone's shooting command, an image containing the display screen of the target large screen is acquired, and embedded hidden code identifiers are extracted and identified in the image. Based on the transformation relationship between the smartphone's camera coordinate system and the world coordinate system, the projection position of the smartphone's camera optical center on the target large screen plane is calculated, and a cursor icon is displayed at the projection position. The pointing position of the smartphone on the large screen is determined by the movement trajectory of the cursor icon. In response to the smartphone pointing at the large screen, the cursor icon moves in real time according to the determined pointing position, and the corresponding selection or control operation is performed on the large screen in response to the user's selection or control command based on the cursor icon.

2. The method for positioning large-screen smartphones based on coded images as described in claim 1, characterized in that, Hidden code identifiers are embedded in several low-texture, color-flat areas of the original display image. Based on the RGB channel cyclic shift, three hidden code embedding frames F1, F2, and F3 are generated, and the original image compensation frame F4 is generated, forming a four-frame cyclic sequence at a predetermined frequency. This cyclic sequence is used to control the alternating playback sequence of the hidden code frame, compensation frame, and original frame according to the refresh rate of the set frequency, ensuring that the hidden code is invisible to the human eye, and the human eye only sees the continuous and complete original display image.

3. The method for large-screen positioning based on coded information and smartphones as described in claim 2, characterized in that, The process of generating three hidden code embedding frames (F1, F2, and F3) based on RGB channel cyclic shifting, and then generating the original image compensation frame (F4), includes: using the B channel of the original displayed image as the background; displaying the hidden code white pixels using the R channel; and displaying the hidden code black pixels using the G channel, thus generating a blue background hidden code frame. frame; The R channel of the original displayed image is used as the background; the hidden white pixels are displayed using the G channel; and the hidden black pixels are displayed using the B channel, generating a red background hidden frame. frame; The G channel of the original displayed image is used as the background; the hidden white pixels are displayed using the B channel; and the hidden black pixels are displayed using the R channel, generating a green background hidden frame. frame; It directly displays the original color image or original grayscale image, without embedding any hidden identifiers, and is only used for image brightness compensation, color restoration, and visual comfort enhancement, generating original image compensation frames. frame.

4. The method for large-screen positioning based on coded information and smartphones as described in claim 3, characterized in that, The hidden identifier is in , , In the three frames, the hidden code is embedded into the image through an RGB channel cyclic shift strategy, creating a significant tonal difference between the hidden code and the background. The frames do not carry hidden code information and are used to offset the color shift and brightness decay of the first three frames. The four frames are output continuously in a fixed order. The human eye cannot perceive the content of a single frame and can only observe the original image after fusion without any markings, thus achieving complete visual hiding of the hidden code.

5. The method for large-screen positioning based on coded information and smartphones as described in claim 1, characterized in that, The process of acquiring an image containing the target large screen display in response to a smartphone's shooting command, and extracting and recognizing the pre-hidden embedded steganographic identifiers in the image includes: converting the steganographic identifier image captured by the smartphone from the RGB color space to the HSV color space, extracting the hue H channel, normalizing the value of the hue H channel, generating three hue offset copies, converting the three copies into grayscale images and performing binarization and morphological denoising, simultaneously performing steganographic identifier recognition on the three hue offset copies, counting the number of successful recognitions for each copy, selecting the copy with the most recognitions as the final valid steganographic identifier data, recognizing and decoding each complete steganographic identifier structure contained in the image, obtaining the unique ID code information corresponding to each steganographic identifier structure and its spatial pose transformation data relative to the camera coordinate system.

6. The method for large-screen positioning based on coded information and smartphones as described in claim 5, characterized in that, The process of extracting the hue H channel, normalizing the values ​​of the hue H channel, and generating three hue offset copies includes: normalizing the values ​​of the hue H channel to the [0, 1] interval to eliminate the numerical ambiguity of red hues at 0° and 360°. For the normalized H-channel image, three independent tone-shifted copies are generated, where the H-channel of the first copy remains unchanged, denoted as . The second copy's H channel is offset by a certain value M, denoted as... ,like Then update The H channel of the third copy is offset by 2M, denoted as... ,like Then update .

7. The method for large-screen positioning based on coded information and smartphones as described in claim 5, characterized in that, The process of counting the number of successful identifications for each copy and selecting the copy with the most identifications as the final valid hidden code data includes: performing hidden code detection and decoding in parallel on the binarized candidate image of each copy, that is, performing corner detection, edge extraction and ID decoding on each copy; counting the number of hidden codes successfully identified for each copy; selecting the copy with the most identifications as the final valid hidden code data, and if multiple copies have the same number of identifications, selecting the result with the highest confidence.

8. The method for large-screen positioning based on coded information and smartphones as described in claim 1, characterized in that, The process of calculating the projection position of the smartphone camera's optical center on the target large screen plane based on the transformation relationship between the smartphone's camera coordinate system and the world coordinate system includes: establishing a world coordinate system with the endpoint of the large screen as the origin, where the X-axis of the world coordinate system is the horizontal direction of the screen, the Y-axis is the vertical direction of the screen, and the Z-axis is perpendicular to the screen plane and points outward; calculating a set of candidate projection points of the smartphone's optical center on the screen plane based on the spatial pose transformation data of the hidden identifier relative to the camera coordinate system, and the transformation relationship between the smartphone's camera coordinate system and the world coordinate system; then performing interpolation calculations on the candidate projection points to finally obtain the precise coordinates of the smartphone's optical center mapped on the screen.

9. The method for large-screen positioning based on coded information and smartphones as described in claim 1, characterized in that, The hidden identifier is a QR code.

10. A large-screen positioning system based on coded information and smartphones, characterized in that, include: The display-end hidden code embedding module is used to embed hidden code identifiers into the original display content of the large screen, generate hidden code identifier frames, difference compensation frames, and RGB channel cyclic shift frames, and construct a cyclic frame sequence that meets the requirements of flicker fusion, so that the hidden code identifiers are invisible to the human eye and only the original display content is displayed. The mobile phone pointing and positioning calculation module is used to respond to the shooting command of the smartphone, acquire an image containing the display screen of the target large screen, extract and identify the embedded hidden code identifiers in the image, calculate the projection position of the optical center of the smartphone camera on the plane of the target large screen according to the transformation relationship between the smartphone's camera coordinate system and the world coordinate system, display a cursor icon at the projection position, and determine the pointing position of the smartphone on the large screen by the movement trajectory of the cursor icon. The human-computer interaction module is used to respond to the operation of the smartphone pointing at the large screen. According to the determined pointing position, it controls the cursor icon to move in real time and responds to the user's selection or control command based on the cursor icon to perform the corresponding selection or control operation on the large screen.