Dynamic recommendation method and device based on portrait recognition

By aligning facial recognition technology with signal feature data from a distributed group of devices, and combining user historical preferences and behavioral priors, proactive streaming commands are generated. This solves the problem of discontinuous cross-device viewing in traditional recommendation systems, enabling seamless cross-device streaming of personalized content and improving recommendation accuracy.

CN122053916APending Publication Date: 2026-05-15WUHAN FENGXING ONLINE TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUHAN FENGXING ONLINE TECH CO LTD
Filing Date
2025-12-24
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Traditional recommendation systems cannot accurately capture users' real-time viewing needs, resulting in a disjointed viewing experience across devices, requiring users to manually search or relocate.

Method used

By acquiring users' visual feature data through facial recognition technology and aligning it with the signal feature data of a distributed group of devices, a cross-device observation feature stream is generated. Combining users' historical preferences and behavioral priors, the posterior probability of identity and migration intent are calculated, and an active transfer command is output to achieve seamless cross-device content transfer.

Benefits of technology

It enables the cross-device flow of personalized content, improves the accuracy and continuity of recommendations, ensures seamless content switching when users reach their target room, and enhances the consistency and personalization of the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122053916A_ABST
    Figure CN122053916A_ABST
Patent Text Reader

Abstract

The invention provides a dynamic recommendation method and device based on portrait recognition, and relates to the technical field of data circulation, and the method comprises the steps: collecting a user portrait video stream to extract visual features, and aligning the visual features with signal features collected by a distributed device group to generate a cross-device observation feature stream; on the basis, target recommendation content is generated in combination with historical preferences of the user, and the identity posterior probability is calculated through logarithm likelihood ratio weighted fusion. And the migration intention of the user is inferred by combining the movement direction, the passage accessibility and the behavior prior, and the room arrival probability distribution and the predicted arrival time are obtained. And when the identity posterior probability and the room arrival probability meet a threshold value and the predicted arrival time is less than the equipment preparation time delay, generating an active circulation command in advance, and triggering seamless content circulation when the user arrives at the boundary of the target room. According to the method, cross-device circulation of personalized contents can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of data flow, specifically to a dynamic recommendation method and apparatus based on facial recognition. Background Technology

[0002] With the increasing prevalence of smart TVs in homes, the viewing experience is gradually shifting from single-device viewing to multi-device interaction, and users are increasingly watching content on different TVs in different rooms. However, traditional recommendation systems mainly rely on users' historical viewing records or popular content lists provided by platforms, failing to accurately capture users' immediate viewing needs and making it difficult to meet the seamless integration between multiple devices. If a user hasn't finished watching a program in the living room and wants to continue watching it on the bedroom TV, they often need to manually search or relocate, resulting in a fragmented experience.

[0003] Against this backdrop, achieving personalized recommendations and cross-device continuation of viewing becomes a key issue. By sensing the current playback status of the TV in real time and combining user identification and viewing preference analysis, the content currently playing can be directly pushed to another TV device, enabling cross-device content flow. This not only solves the problem of continuation of viewing in multi-member, multi-TV scenarios, but also dynamically incorporates the user's real-time viewing context into the recommendation mechanism, thereby improving the accuracy and continuity of recommendations.

[0004] Therefore, a method is needed to enable the cross-device flow of personalized content. Summary of the Invention

[0005] This invention provides a dynamic recommendation method and apparatus based on facial recognition, which enables personalized content to be transferred across devices.

[0006] A first aspect of the present invention provides a dynamic recommendation method based on facial recognition, the method comprising: The user's visual feature data is obtained by processing the user's portrait video stream and aligning it with the signal feature data shared by the distributed device group to output a cross-device observation feature stream; Based on the visual feature data and combined with the user's historical preference data, target recommended content is generated; Using the cross-device observation feature stream as input, the posterior probability of the user's identity is calculated based on log-likelihood ratio weighted fusion; Based on the posterior probability of identity and the cross-device observation feature stream, the user's migration intention is calculated based on the direction of movement, accessibility, and behavioral prior, and the room arrival probability distribution and estimated arrival time of different candidate rooms are output. When both the identity posterior probability and the room arrival probability of the target room meet the trigger threshold, and the estimated arrival time of the target room is less than the device preparation delay of the target room, an active transfer command is output. According to the active flow command, when the user is detected to have reached the boundary of the target room, the target recommended content is flowed to the device in the target room.

[0007] Based on the above technical solutions, preferably, the step of generating target recommendation content based on the visual feature data and combined with user historical preference data specifically includes: After the visual feature data is standardized and mapped and aligned, it is spatially aligned with the historical embedding vector corresponding to the user's historical preference data to obtain the visual embedding and the historical embedding. The visual embedding and the historical embedding are dynamically adjusted based on the intensity and consistency of the visual features to determine the weights of short-term interests and long-term preferences, and a personalized state vector is output. By combining the personalized state vector with the age classification information and scene context information determined based on the visual feature data, a constraint mask is generated to filter out the content and obtain candidate content. The personalized state vector is matched and scored with the candidate content vector, and the original score set of the candidate content is generated by combining the session context vector and the platform quality factor. Based on the original score set, probability normalization is performed under the constraint mask to obtain target recommended content that meets the display conditions.

[0008] Based on the above technical solutions, preferably, the step of calculating the user's posterior probability of identity using the cross-device observation feature stream as input and based on log-likelihood ratio weighted fusion specifically includes: The cross-device observation feature stream is processed for time and space alignment using a unified clock and a unified coordinate reference, and the output is a feature tensor sequence organized by time window. For the visual features and signal features in the feature tensor sequence, feature reliability weights are generated respectively, and the feature reliability weights are adaptively updated based on consistency score, drift score and signal-to-noise ratio score; Based on the feature tensor sequence and the feature reliability weights, modal likelihood models are established under user conditions and non-user conditions, respectively, and the modal level log-likelihood ratio sequence is output. The modal-level log-likelihood ratio sequence and the feature reliability weights are weighted and accumulated to form an immediate posterior probability of identity. The real-time posterior probability of identity is smoothed with the posterior probability of identity in the previous time window, and the smoothed posterior probability of identity and the updated temporal fusion coefficient are output. When the smoothed posterior identity probability is lower than the confidence threshold, the modality selection and re-estimation process is triggered to suppress modalities with significant drift and output the final posterior identity probability.

[0009] Based on the above technical solutions, preferably, the step of calculating the user's migration intention based on the identity posterior probability and the cross-device observation feature stream, and outputting the room arrival probability distribution and estimated arrival time for different rooms, specifically includes: The cross-device observation feature stream is divided into sliding time windows under a unified clock and a unified coordinate reference. The skeleton key point trajectory and optical flow direction are generated from the visual features. The inertial measurement step frequency, ultra-wideband relative displacement increment and wireless channel state change are generated from the signal features. The velocity vector and the unitized motion direction vector are fused together. The velocity vector and room-level topology map are used as inputs, and combined with the functional data of different candidate rooms, a accessibility score is generated. Based on the room history entry records and the current time period label of the unified clock index as input, a behavioral prior with time decay is generated. The room arrival probability distribution of each candidate room is calculated using the unitized motion direction vector, the accessibility score, and the behavioral prior, and the estimated arrival time is calculated by combining the velocity vector and the path length of the room-level topology map.

[0010] Based on the above technical solutions, preferably, when both the identity posterior probability and the room arrival probability of the target room meet the trigger threshold, and the estimated arrival time of the target room is less than the device preparation delay of the target room, an active transfer command is output, specifically including: The posterior probability of the identity and the probability of arriving at the room are written into the decision context under a unified clock, and the target room index and the corresponding arrival time window are generated. The target room index is invoked as input to decompose the device preparation delay, which includes device discovery and authentication delay, media session pre-establishment delay, start-up buffer filling delay, and timeline alignment delay. After comparing the device preparation delay with the estimated arrival time, if the estimated arrival time is less than the device preparation delay and the identity posterior probability and the room arrival probability meet the trigger threshold, an active transfer command is generated.

[0011] Based on the above technical solutions, preferably, the device that, according to the active transfer command, transfers the target recommended content to the target room when it detects that the user has reached the boundary of the target room, further includes: The active flow command triggers boundary event detection in the door frame area. The boundary event detection generates arrival boundary markers and arrival timestamps based on infrared door magnetic triggers, ultra-wideband near-field threshold events and camera door frame intrusion events, and writes the arrival boundary markers and arrival timestamps into the decision context. When the target room device receives the active streaming command and holds a media session identifier, it sends mute and frame freeze commands through the distributed soft bus, and simultaneously sends unmute and playback unfreeze commands to the target room device. While the target room device completes playback unfreezing according to the timeline markers, the current room device is controlled to stop audio output and freeze the rendering frame.

[0012] Based on the above technical solutions, preferably, the step of processing the user's facial video stream to obtain the user's visual feature data, aligning it with the signal feature data shared by the distributed device group, and outputting a cross-device observation feature stream specifically includes: The distributed soft bus is used to complete the device discovery, authentication and unified clock synchronization of the distributed device group, and to establish a distributed clock and unified coordinate reference. The human image video stream is captured by the camera of the master device in the distributed device group, and video frames are extracted by a key frame extraction algorithm. The video frames are input into a deep learning model to generate the user's portrait re-identification embedding, skeleton key point motion vectors and gait embedding, forming the visual feature data; The wireless channel status information, Bluetooth signal strength trajectory, ultra-wideband ranging data and inertial measurement data are collected by other devices in the distributed device group other than the main device to form the signal feature data stream; The visual feature data and the signal feature data are timestamped using the distributed clock, and the spatial location data is normalized based on the unified coordinate reference to output the cross-device observation feature stream.

[0013] In a second aspect of the invention, a dynamic recommendation device based on facial recognition is provided. The device is used to execute a dynamic recommendation method based on facial recognition as described in any of the above embodiments. The device includes an acquisition module, a processing module, and an output module, wherein: The acquisition module is used to process the user's visual feature data obtained by collecting the user's portrait video stream, align it with the signal feature data shared by the distributed device group, and output a cross-device observation feature stream. The processing module is used to generate target recommended content based on the visual feature data and combined with the user's historical preference data; The processing module is used to calculate the posterior probability of the user's identity based on the cross-device observation feature stream as input and weighted fusion of log-likelihood ratio. The processing module is used to calculate the user's migration intention based on the identity posterior probability and the cross-device observation feature stream, and based on the direction of movement, accessibility, and behavioral priors, and output the room arrival probability distribution and estimated arrival time for different candidate rooms; The processing module is configured to output an active transfer command when both the identity posterior probability and the room arrival probability of the target room meet the trigger threshold, and the estimated arrival time of the target room is less than the device preparation delay of the target room. The output module is used to transfer the target recommended content to the device in the target room when the user is detected to have reached the boundary of the target room, according to the active transfer command.

[0014] In a third aspect of the invention, an electronic device is provided, including a processor, a memory, a user interface, and a network interface, wherein the memory is used to store instructions, the user interface and the network interface are both used to communicate with other devices, and the processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any of the preceding embodiments.

[0015] In a fourth aspect, the present invention provides a computer-readable storage medium storing instructions that, when executed, perform the method as described in any of the preceding claims.

[0016] In summary, one or more technical solutions provided in the embodiments of the present invention have at least the following technical effects or advantages: 1. This invention aligns the visual feature data obtained from processing the user's portrait video stream with the signal feature data shared by a distributed group of devices to form a cross-device observation feature stream. Based on this, it generates target recommended content by combining the user's historical preferences, and calculates the posterior probability of identity using log-likelihood ratio weighted fusion. Then, it infers the user's migration intention and expected arrival time through movement direction, accessibility, and behavioral priors. Thus, it issues an active transfer command in advance when the triggering conditions are met, and seamlessly switches the target recommended content to the target device when the user is detected to have reached the boundary of the target room. This achieves comprehensive perception and prediction of user identity, real-time behavior, and viewing context, thereby enabling the cross-device transfer of personalized content while ensuring continuity and accuracy.

[0017] 2. By aligning and integrating visual feature data with users' historical preference data, and combining personalized status, age classification, and scene context, dynamic filtering and matching scoring are achieved, thereby generating recommended content that matches users' immediate interests and long-term preferences, improving the accuracy and personalization of recommendations.

[0018] 3. By utilizing cross-device observation feature streams and combining feature reliability weights with log-likelihood ratio weighted fusion, robust aggregation of multimodal evidence is achieved. Furthermore, noise and drift are suppressed through temporal smoothing and modality selection mechanisms to ensure the accuracy and continuity of posterior identity probabilities.

[0019] 4. By jointly modeling movement direction, accessibility, and behavioral priors, the probability distribution of room arrival and the estimated arrival time are calculated, thereby enabling accurate prediction of user migration intentions, providing a basis for proactive cross-room movement and improving the reliability of predictions.

[0020] 5. Under the premise of credible identity and clear migration intention, introduce a device preparation delay assessment mechanism. By comparing the expected arrival time with the preparation delay, ensure that the switching timing is reasonable, thereby generating an active transfer command to achieve pre-warm-up and avoid waiting after the user arrives.

[0021] 6. When the user reaches the boundary of the target room, multimodal event detection is triggered, and control commands are synchronously issued using a distributed soft bus. While the target room device is unfrozen and playing, the current device is frozen, achieving seamless switching of audio and video frames and ensuring the continuity and naturalness of content flow. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating a dynamic recommendation method based on facial recognition disclosed in an embodiment of the present invention; Figure 2 This is a schematic diagram of a dynamic recommendation device based on facial recognition disclosed in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention.

[0023] Explanation of reference numerals in the attached drawings: 201, acquisition module; 202, processing module; 203, output module; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Implementation

[0024] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0025] In the description of the embodiments of the present invention, words such as "for example" or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design described as "for example" or "for instance" in the embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of words such as "for example" or "for instance" is intended to present the relevant concepts in a specific manner.

[0026] In the description of the embodiments of the present invention, the term "multiple" means two or more. For example, multiple systems means two or more systems, and multiple screen terminals means two or more screen terminals. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, a feature defined with "first" or "second" may explicitly or implicitly include one or more of that feature. The terms "comprising," "including," "having," and variations thereof all mean "including but not limited to," unless otherwise specifically emphasized.

[0027] As smart TVs become increasingly common in homes, the demand for users to watch content across rooms and multiple devices is growing. However, traditional recommendation systems rely solely on historical records or popular lists, failing to meet the immediacy and continuity of cross-device viewing, forcing users to manually search and creating a fragmented experience. To address this issue, a method is needed that combines user identification and viewing preference analysis. Based on real-time awareness of the current playback status, content can be proactively pushed to another TV, enabling personalized content flow across devices and thus improving the accuracy of recommendations and the continuity of the viewing experience.

[0028] This embodiment discloses a dynamic recommendation method based on facial recognition, referring to... Figure 1 This includes the following steps S110-S160: S110 collects the user's portrait video stream, processes the resulting visual feature data, aligns it with the signal feature data shared by the distributed device group, and outputs a cross-device observation feature stream.

[0029] This invention discloses a dynamic recommendation method based on facial recognition, applied to a server. The server includes, but is not limited to, electronic devices such as mobile phones, tablets, wearable devices, and PCs (Personal Computers), and can also be a backend server running the dynamic recommendation method based on facial recognition. The server can be implemented using a standalone server or a server cluster composed of multiple servers.

[0030] In one possible implementation, the user's visual feature data, obtained by processing the user's facial video stream, is aligned with the signal feature data shared by the distributed device group to output a cross-device observation feature stream. Specifically, this includes: completing device discovery, authentication, and unified clock synchronization for the distributed device group via a distributed soft bus, establishing a distributed clock and a unified coordinate reference; capturing the facial video stream using the camera of the master device within the distributed device group, and extracting video frames using a keyframe extraction algorithm; inputting the video frames into a deep learning model to generate the user's facial re-identification embedding, skeleton keypoint motion vectors, and gait embedding, forming visual feature data; collecting wireless channel status information, Bluetooth signal strength trajectory, ultra-wideband ranging data, and inertial measurement data from other devices within the distributed device group (excluding the master device), forming a signal feature data stream; aligning the visual feature data and signal feature data with timestamps using a distributed clock, and normalizing the spatial location data based on the unified coordinate reference, outputting a cross-device observation feature stream.

[0031] Specifically, the distributed soft bus, as a communication middleware between devices, provides low-latency discovery and highly reliable authentication mechanisms across terminals. Upon initial device access, the master device broadcasts a discovery request to the local area network (LAN). Other devices respond and return device identity information, including a unique device identifier and a description of device capabilities. The authentication process is achieved through security certificate verification and key exchange, ensuring trustworthy interaction between devices. After discovery and authentication, the distributed soft bus performs unified clock synchronization. This synchronization is achieved through timestamp message round-trip delay estimation and drift compensation mechanisms, ensuring all devices collect and process data on the same time base. A unified coordinate reference is obtained through spatial topology modeling between devices, specifically employing a coordinate alignment method based on signal ranging and fixed positioning points to guarantee a unified representation of spatial data across devices.

[0032] The main device is typically a television with the highest computing power and optimal network bandwidth, equipped with a high-resolution camera for real-time image capture of the user. The captured video stream is decoded and then input into a keyframe extraction algorithm. This algorithm comprehensively evaluates image sharpness, face detection stability, and inter-frame change rate to select video frames containing clear facial information of the user. This processing ensures the quality stability of the input images for subsequent feature extraction, avoiding performance degradation due to blurring or uneven lighting.

[0033] The deep learning model comprises several sub-network modules: Human face re-identification embedding is jointly implemented using a convolutional neural network and an attention mechanism, with the output vector uniquely representing user identity features. Skeletal keypoint motion vectors are generated by a pose estimation network, which detects the positions of key joints in the human body and outputs motion vectors through time-series modeling to characterize the user's dynamic behavior. Gait embedding is generated by a fusion model based on temporal convolution and recurrent neural networks. This model utilizes changes in the overall shape and movement rhythm of the human body in consecutive frames to extract unique gait features of the user during walking and movement. These features collectively constitute visual feature data, forming a complete visual representation of the user.

[0034] Signal feature acquisition is achieved through multimodal sensors. Wireless channel state information is obtained through a channel state information estimation algorithm at the receiver, reflecting the impact of spatial position changes on the wireless propagation path. Bluetooth signal strength trajectory is formed by continuously acquiring the received power value of Bluetooth Low Energy signals, with power fluctuating with distance and obstacles. Ultra-wideband ranging data is acquired by the ultra-wideband transceiver module, utilizing the time-of-flight principle to achieve high-precision positioning. Inertial measurement data is provided by the inertial measurement unit, including the outputs of a three-axis accelerometer and a three-axis gyroscope, used to capture acceleration and angular velocity information during user movement. The above multi-source data are fused to form a signal feature data stream, characterizing the user's state from three aspects: wireless signal propagation, short-range communication, and inertial motion.

[0035] Timestamp alignment is achieved through a globally distributed clock, ensuring that visual and signal features originate from the same time window and avoiding temporal misalignment. Spatial location data normalization uses a unified coordinate reference, employing linear or rigid transformations to map location data from different device coordinate systems to the global coordinate system. This process compensates for device installation angles and ranging offsets to ensure data consistency across spatial dimensions. After time alignment and spatial normalization, a cross-device observation feature stream is output. This stream integrates visual and signal features, providing a unified input for subsequent identity posterior probability calculation and migration intent prediction.

[0036] S120 generates target recommended content based on visual feature data and combined with user historical preference data.

[0037] In one possible implementation, target recommended content is generated based on visual feature data and user historical preference data. Specifically, this includes: standardizing and mapping the visual feature data, then spatially aligning it with the historical embedding vectors corresponding to the user's historical preference data to obtain visual embeddings and historical embeddings; dynamically adjusting the weights of short-term interests and long-term preferences based on visual feature intensity and consistency, outputting a personalized state vector; combining the personalized state vector with age grading information and scene context information determined based on the visual feature data to generate a constraint mask to filter out content and obtain candidate content; matching and scoring the personalized state vector with the candidate content vector, and combining the session context vector with the platform quality factor to generate an original score set for the candidate content; and performing probability normalization based on the original score set under the constraint mask to obtain target recommended content that meets the display conditions.

[0038] Specifically, the visual feature data, after standardization and mapping alignment, is spatially aligned with the historical embedding vectors corresponding to the user's historical preference data. First, the visual feature data and historical embedding vectors are scaled and calibrated to eliminate numerical differences and statistical biases. Then, a linear mapping matrix projects both types of data into a unified embedding space, outputting the visual embedding and the historical embedding. The formula is as follows:

[0039]

[0040] in, Embed vectors for misaligned visual feature data. For unaligned historical embedding vectors, and It is a linear mapping matrix. and The sample mean vector and This is the standard deviation vector of the samples. This step ensures that visual data and historical data are comparable within a unified semantic space.

[0041] The weights of short-term interests and long-term preferences are dynamically adjusted between visual embedding and historical embedding based on visual feature strength and consistency. Visual feature strength is jointly determined by keyframe confidence, facial expression stability, and pose stability. Consistency is calculated using the rate of change of the embedding direction within a continuous time window. Finally, a gating mechanism is used to determine the fusion ratio and output a personalized state vector. The formula is as follows:

[0042] in, This represents the short-term interest weight, with values ​​ranging from zero to one. For logical functions, For the weight vector, For bias, This is a personalized state vector. This step ensures that short-term interests and long-term preferences are balanced under dynamic weights.

[0043] A constraint mask is generated by combining personalized state vectors with age grading information determined from visual feature data and scene context information. Age grading information is determined by facial age estimation results and user monitoring parameters. Scene context information includes room labels, time period labels, and interaction intensity labels. The constraint mask is used to filter out non-compliant or mismatched content, outputting a set of candidate content. The formula is as follows:

[0044] in, Availability marker for candidate content. Content is categorized and tagged. The highest allowed rating, For scene tags, This is a set of playable scenarios. This step ensures that the recommended content matches the user's attributes and usage scenario.

[0045] The personalized state vector is matched and scored with the candidate content vector. Then, the session context vector and platform quality factors are combined to generate the original score set for the candidate content. The session context vector represents the semantic direction and continuity of the currently playing content. The platform quality factors include content completion rate, freshness, and compliance priority. The formula is as follows:

[0046] in, The original score for the candidate content. For content semantic embedding, For the session context vector, For platform quality factors, , , These are weighting coefficients. This step ensures that the recommendation results take into account user preferences, continuity, and platform constraints.

[0047] Based on the original score set, probability normalization is performed under the constraint mask to output target recommended content that meets the display conditions. A temperature parameter is introduced during the normalization process to adjust the balance between exploration and utilization. The formula is as follows:

[0048]

[0049] in, The probability of selecting candidate content. For temperature parameters, To constrain mask markings, This is the original score. This step ensures that high-scoring content is prioritized in the recommendation results, while also preserving a certain degree of diversity, ultimately outputting the target recommended content list.

[0050] S130 takes cross-device observation feature streams as input and calculates the posterior probability of a user's identity based on log-likelihood ratio weighted fusion.

[0051] In one possible implementation, the user's posterior probability is calculated based on a weighted fusion of log-likelihood ratios, using a cross-device observation feature stream as input. Specifically, this includes: time-aligning and spatial-aligning the cross-device observation feature stream using a unified clock and coordinate reference, outputting a feature tensor sequence organized by time windows; generating feature reliability weights for visual and signal features in the feature tensor sequence, with the feature reliability weights adaptively updated based on consistency scores, drift scores, and signal-to-noise ratio scores; establishing modal likelihood models under user and non-user conditions based on the feature tensor sequence and feature reliability weights, outputting a modal-level log-likelihood ratio sequence; performing weighted evidence accumulation on the modal-level log-likelihood ratio sequence and feature reliability weights and mapping it to an immediate posterior probability of identity; performing temporal smoothing on the immediate posterior probability of identity and the posterior probability of identity in the previous time window, outputting the smoothed posterior probability of identity and the updated temporal fusion coefficients; triggering a modal selection and re-estimation process when the smoothed posterior probability of identity is lower than a confidence threshold, suppressing modes with significant drift, and outputting the final posterior probability of identity.

[0052] Specifically, the cross-device observation feature streams are time-aligned and spatially aligned using a unified clock and a unified coordinate reference. First, a global timestamp is assigned to each observation based on the unified clock, and an equally spaced time series is formed through resampling and interpolation. Then, using the unified coordinate reference, the coordinates of each device are transformed to the global coordinate system, and installation attitude and ranging offsets are compensated for, resulting in a three-dimensional feature tensor sequence organized by time windows, with channels stacked according to visual and signal features. The formula is as follows:

[0053]

[0054] in, This is the position vector in the global coordinate system; This is the position vector in the device coordinate system; It is a rotation matrix; This is the translation vector. The principle behind this formula is rigid body transformation to unify the spatial reference.

[0055] Feature reliability weights are generated for both visual and signal features in the feature tensor sequence. For each modality, a consistency score, drift score, and signal-to-noise ratio score are calculated and normalized to a range of zero to one. These scores are then fused using a weighted average to form modality-level feature reliability weights, which are adaptively updated according to a time window. The formula is as follows:

[0056]

[0057] in, For the first Modal characteristic reliability weights; For consistency scoring; For drift rating; Score the signal-to-noise ratio; , , These are non-negative fusion coefficients. The calculation principle is to fuse multiple indicators into probabilistic weights using a soft-maximization method.

[0058] Modal likelihood models under user-conditional and non-user-conditional features are established based on the feature tensor sequence and feature reliability weights, respectively. Visual features are estimated using embedding-based class-conditional density estimation, while signal features are estimated using kernel density estimation under trajectory conditions, with bandwidth adaptation based on context. This outputs a modal-level log-likelihood ratio sequence. The formula is as follows:

[0059]

[0060] in, For the first The log-likelihood ratio of the modes; This represents the observational characteristics of this mode within the current time window; Likelihood under user conditions; This represents the likelihood under non-user conditions. The calculation principle is to measure the strength of evidence using log-odds.

[0061] A weighted evidence accumulation is performed on the modal-level log-likelihood ratio sequence and feature reliability weights, and mapped to the immediate posterior probability of identity. The calculation process involves evidence summation in the log-odds domain, followed by a logistic function to obtain the probability. The formula is as follows:

[0062]

[0063]

[0064] in, Logarithmic odds; For the user's prior probability; For feature reliability weights; The modal-level log-likelihood ratio; This represents the immediate posterior probability of identity. The calculation principle is based on the additivity of multi-source evidence in the logarithmic field, and then the mapping from probability to likelihood is completed by a logistic function.

[0065] The immediate posterior probability of identity is time-series smoothed compared to the posterior probability of identity in the previous time window to reduce noise impact and maintain response speed. Exponential smoothing is used, as shown in the following formula:

[0066]

[0067] in, This represents the smoothed posterior probability of identity for the current time window. For immediate posterior probability; This represents the smoothed probability of the previous time window; The time-series fusion coefficient ranges from zero to one. The calculation principle is a convex combination of new evidence and historical states to suppress short-term jitter.

[0068] When the smoothed posterior identity probability falls below a confidence threshold, a modality selection and re-estimation process is triggered. Gating conditions are set based on modality-level feature reliability weights and drift scores to mask low-confidence modalities, and the posterior identity probability is recalculated using the remaining modalities. The formula is as follows:

[0069]

[0070]

[0071] in, For gating; This is the lower limit of the weight; This is the upper limit of drift; The logarithmic probability after gating; This represents the posterior probability of the final identity. The calculation principle is to eliminate unstable modes and retain stable evidence to achieve robust estimation.

[0072] S140 calculates the user's migration intention based on the posterior probability of identity and the cross-device observation feature flow, and calculates the user's migration intention based on the direction of movement, accessibility, and behavioral priors. It then outputs the room arrival probability distribution and estimated arrival time for different candidate rooms.

[0073] In one possible implementation, based on the posterior probability of identity and cross-device observation feature flow, the user's migration intention is calculated based on motion direction, accessibility, and behavioral priors. The result is the output of room arrival probability distributions and estimated arrival times for different rooms. Specifically, this includes: dividing the cross-device observation feature flow into sliding time windows under a unified clock and coordinate reference; generating skeleton keypoint trajectories and optical flow directions from visual features; generating inertial measurement step frequency, ultra-wideband relative displacement increment, and wireless channel state change from signal features; fusing these to obtain velocity vectors and unitized motion direction vectors; using the velocity vectors and room-level topology map as input, and combining functional data from different candidate rooms, generating accessibility scores; using historical room entry records with a unified clock index and the current time period label as input, generating behavioral priors with time decay; calculating the room arrival probability distribution for each candidate room using the unitized motion direction vectors, accessibility scores, and behavioral priors; and calculating the estimated arrival time by combining the velocity vectors and path lengths from the room-level topology map.

[0074] Specifically, the cross-device observation feature stream is divided into sliding time windows under a unified clock and a unified coordinate reference. The unified clock ensures that the sampling data from each device are aligned on the same time base, and the unified coordinate reference ensures that spatial location data is uniformly represented in the global coordinate system. The sliding time window uses a fixed-length and step-size window division method to divide continuous observation data into multiple time segments, so as to extract stable features in the local time domain. Within each time window, the skeleton keypoint trajectory and optical flow direction are extracted from visual features. The skeleton keypoint trajectory consists of the human joint position sequence output by the attitude estimation algorithm, and the optical flow direction is calculated from the pixel motion vectors between adjacent frames, used to characterize the user's movement trend. Inertial measurement step frequency, ultra-wideband relative displacement increment, and wireless channel state change are extracted from signal features. The inertial measurement step frequency is obtained through spectral analysis of the acceleration signal, the ultra-wideband relative displacement increment is calculated through time-of-flight difference, and the wireless channel state change is obtained through statistical analysis of the amplitude and phase changes of channel state information. These visual features and signal features are fused within a time window to obtain a velocity vector and a unitized motion direction vector. The velocity vector represents the magnitude and direction of the user's motion, while the unitized motion direction vector is used to highlight the directional attribute.

[0075] Using velocity vectors and room-level topology maps as input, and combining functional data from different candidate rooms, a accessibility score is generated. The room-level topology map is a graph structure constructed from the access relationships between rooms within the room, containing room nodes and doorways or passageways as edges. Functional data includes door opening / closing status, path congestion, and room availability labels. The accessibility score comprehensively considers three aspects: path connectivity, path smoothness, and functional constraints. For example, when a room's door is closed, its accessibility score decreases; when there is congestion or obstacles on the path, the path smoothness score decreases. Finally, a weighted fusion is used to obtain the accessibility score for each candidate room, reflecting the likelihood that a user will actually reach that room in the physical environment.

[0076] Using historical room entry records with a unified clock index and the current time period label as input, a behavioral prior with time decay is generated. Historical room entry records are obtained by statistically analyzing the frequency of users entering different rooms over a past period. The current time period label reflects the specific time context of the user, such as morning, noon, or evening. The behavioral prior uses a time decay function to weight historical entry records, assigning higher weights to more recent entries and gradually decreasing weights to more distant entries. This process ensures that recent user behavior patterns have a greater impact on prediction, while the influence of more distant behaviors gradually weakens, resulting in a dynamically updated behavioral prior probability distribution.

[0077] The arrival probability distribution of each candidate room is calculated using a unitized motion direction vector, accessibility score, and behavioral prior. The estimated arrival time is then calculated by combining the velocity vector and the path length from the room-level topology map. The room arrival probability distribution is jointly mapped to a normalized probability value using a soft maximum function, where directional consistency is characterized by the similarity of the angle between the unitized motion direction vector and the room entrance direction vector. The estimated arrival time is obtained by dividing the path length by the magnitude of the velocity vector, while incorporating an environmental correction factor to compensate for dynamic obstacles or congestion on the path. The final output includes the arrival probability value and corresponding estimated arrival time for each candidate room, providing a reliable basis for subsequent proactive flow determination and recommended content migration.

[0078] S150, when the posterior probability of identity and the room arrival probability of the target room both meet the trigger threshold, and the expected arrival time of the target room is less than the equipment preparation delay of the target room, output an active transfer command.

[0079] In one possible implementation, when both the identity posterior probability and the room arrival probability of the target room meet the trigger threshold, and the expected arrival time of the target room is less than the device preparation delay of the target room, an active transfer command is output. Specifically, this includes: writing the identity posterior probability and the room arrival probability into the decision context under a unified clock, and generating a target room index and a corresponding arrival time window; using the target room index as input to decompose the device preparation delay, which includes device discovery and authentication delay, media session pre-establishment delay, start-up buffer filling delay, and timeline alignment delay; and after comparing the device preparation delay with the expected arrival time, when the expected arrival time is less than the device preparation delay and the identity posterior probability and the room arrival probability meet the trigger threshold, an active transfer command is generated.

[0080] Specifically, the posterior probability of identity and the probability of room arrival are written into the decision context, globally timestamped according to a unified clock, and a target room index and arrival time window are generated. The decision context record fields include the posterior probability of identity, the probability of room arrival, the target room index, the start and end dates of the arrival time window, the session identifier, the transaction sequence number, and the retry count. The unified clock ensures consistent event sorting across devices. The target room index is given by the room corresponding to the maximum room arrival probability. The arrival time window is composed of a point estimate of the expected arrival time and its uncertainty, with the start and end boundaries formed by adding or subtracting tolerances from the center time. An idempotent flag and signature digest are added upon successful writing for subsequent command orchestration and audit traceability.

[0081] The device preparation latency is decomposed using the target room index as input. It is estimated item by item for device discovery and authentication latency, media session pre-establishment latency, playback buffer filling latency, and timeline alignment latency, and then aggregated to obtain the upper bound and quantile statistics of the preparation latency. Device discovery and authentication latency is the sum of neighborhood discovery time, key negotiation time, and permission verification time; media session pre-establishment latency consists of codec warm-up, decoding pipeline pre-orchestration, and rendering context initialization; playback buffer filling latency is calculated based on bitrate, buffer target level, network throughput, and jitter estimation; timeline alignment latency consists of clock skew measurement, propagation latency estimation, and playback rate fine-tuning convergence time. Each sub-item is estimated using sliding window statistics and device health weighting, and the preparation latency interval, confidence level, and critical path identifier are output and written into the decision context for comparison and judgment.

[0082] After comparing the preparation delay with the estimated arrival time, if the estimated arrival time is less than the preparation delay, and both the identity posterior probability and the room arrival probability are not lower than the trigger threshold, an active transfer command is generated. The command payload includes the target room index, session identifier, timeline identifier, buffer target level, critical path acceleration strategy, timeout and rollback strategy, and transaction sequence number. After the command is signed and timestamped, it is sent to the target room device through the distributed soft bus, and at the same time, the device is set to the "preheating ready" state locally, and status monitoring within the arrival time window is started. If the comparison does not meet the conditions, the status is marked as "observation pending", the decision context is preserved, and a cooldown timer is entered to prevent bouncing and repeated triggering.

[0083] S160, based on the active flow command, when the user is detected to have reached the boundary of the target room, the target recommended content is flowed to the device in the target room.

[0084] In one possible implementation, upon detecting that a user has reached the boundary of the target room, the target recommended content is transferred to the device in the target room according to the active transfer command. Specifically, this includes: triggering boundary event detection of the door frame area based on the active transfer command; generating an arrival boundary marker and arrival timestamp based on an infrared door magnetic trigger, an ultra-wideband near-field threshold event, and a camera door frame intrusion event; and writing the arrival boundary marker and arrival timestamp into the decision context; under the condition that the target room device receives the active transfer command and holds a media session identifier, issuing mute and frame freeze commands through a distributed soft bus, and simultaneously issuing unmute and playback unfreeze commands to the target room device; while the target room device completes playback unfreeze according to the timeline identifier, controlling the current room device to stop audio output and freeze the rendering frame.

[0085] Specifically, boundary event detection in the door frame area is triggered based on the active flow command. Boundary event detection utilizes multimodal sensors installed at the door frame location: an infrared door magnetic trigger detects the door's opening and closing state and generates a trigger signal when the magnetic field changes; an ultra-wideband near-field threshold event measures whether a user has entered the door frame neighborhood based on the time-of-flight difference between the transmitting and receiving nodes and outputs a trigger event when the distance is less than a set threshold; a camera door frame intrusion event identifies the moment a user crosses the door frame boundary using target detection and area intrusion algorithms and generates an intrusion marker. All three types of events are marked with a global timestamp under a unified clock, and an arrival boundary marker and arrival timestamp are generated when any event is triggered. The arrival boundary marker is a binary variable indicating whether a user crossing the boundary has been detected, and the arrival timestamp is the unified clock value at the trigger moment. Both are written into the decision context for subsequent boundary switching synchronization control and audit traceability.

[0086] Once the target room device receives the active flow command and holds a media session identifier, synchronous commands are issued via the distributed soft bus. The current room device issues a mute command and a frame freeze command to itself. The mute command closes the audio output channel, and the frame freeze command stops the rendering and updating of video frames and holds the last frame. Simultaneously, the current room device issues an unmute command and a playback unfreeze command to the target room device. The unmute command opens the audio output channel, and the playback unfreeze command allows aligned video frames in the buffer to enter the rendering process. The distributed soft bus ensures low latency and timing consistency in command transmission between multiple devices and ensures strict alignment of bidirectional actions between the target room device and the current room device through transaction sequence numbers.

[0087] While the target room device completes playback unfreezing according to the timeline marker, the current room device is controlled to execute an audio / video stop. The timeline marker is maintained by a distributed clock calibration mechanism to ensure that the two devices switch at the same frame boundary. After receiving the playback unfreezing command, the target room device switches its media decoding pipeline and rendering pipeline to the active state and starts continuous playback from the buffer frame corresponding to the timeline marker. The current room device then performs a stop operation under the same timeline marker, including turning off the audio output and freezing the screen, so that the user's auditory and visual perception remains continuous without overlap or interruption. This synchronization process uses the timeline marker as a global anchor point to ensure seamless switching of active flow across devices at the frame level, avoiding audiovisual misalignment and perceived latency.

[0088] This embodiment also discloses a dynamic recommendation device based on facial recognition, referring to... Figure 2 The device includes an acquisition module 201, a processing module 202, and an output module 203. It is used to execute any of the above-described dynamic recommendation methods based on facial recognition, wherein: The acquisition module 201 is used to process the user's visual feature data obtained by collecting the user's portrait video stream, align it with the signal feature data shared by the distributed device group, and output a cross-device observation feature stream.

[0089] Processing module 202 is used to generate target recommended content based on visual feature data and combined with user historical preference data.

[0090] Processing module 202 is used to calculate the posterior probability of a user's identity based on a log-likelihood ratio weighted fusion, taking cross-device observation feature streams as input.

[0091] The processing module 202 is used to calculate the user's migration intention based on the identity posterior probability and cross-device observation feature flow, and based on the direction of movement, accessibility and behavioral prior, and output the room arrival probability distribution and expected arrival time of different candidate rooms.

[0092] The processing module 202 is used to output an active transfer command when both the identity posterior probability and the room arrival probability of the target room meet the trigger threshold, and the expected arrival time of the target room is less than the device preparation delay of the target room.

[0093] The output module 203 is used to transfer target recommended content to the target room device when the user is detected to have reached the boundary of the target room, based on the active transfer command.

[0094] In one possible implementation, the processing module 202 is used to spatially align the visual feature data after standardization and mapping alignment with the historical embedding vector corresponding to the user's historical preference data to obtain the visual embedding and the historical embedding.

[0095] The processing module 202 is used to dynamically adjust the weights of short-term interests and long-term preferences based on the visual feature intensity and consistency of the visual embedding and historical embedding, and output a personalized state vector.

[0096] The processing module 202 is used to combine the personalized state vector with age classification information and scene context information determined based on visual feature data to generate a constraint mask to filter out the filtered content and obtain candidate content.

[0097] The processing module 202 is used to match and score the personalized state vector with the candidate content vector, and combine the session context vector with the platform quality factor to generate the original score set of the candidate content.

[0098] The processing module 202 is used to perform probability normalization based on the original rating set under the action of the constraint mask to obtain the target recommended content that meets the display conditions.

[0099] In one possible implementation, the acquisition module 201 is used to perform time and space alignment processing on cross-device observation features through a unified clock and a unified coordinate reference, and output a feature tensor sequence organized by time window.

[0100] The processing module 202 is used to generate feature reliability weights for visual features and signal features in the feature tensor sequence, respectively. The feature reliability weights are adaptively updated based on consistency score, drift score and signal-to-noise ratio score.

[0101] The processing module 202 is used to establish modal likelihood models under user conditions and non-user conditions based on the feature tensor sequence and feature reliability weights, and output the modal level log-likelihood ratio sequence.

[0102] Processing module 202 is used to perform weighted evidence accumulation on modal-level log-likelihood ratio sequences and feature reliability weights and map them to instantaneous posterior probabilities of identity.

[0103] The processing module 202 is used to perform temporal smoothing on the immediate posterior probability of identity and the posterior probability of identity in the previous time window, and output the smoothed posterior probability of identity and the updated temporal fusion coefficient.

[0104] The processing module 202 is used to trigger the modality selection and re-estimation process when the smoothed identity posterior probability is lower than the confidence threshold, suppress modalities with significant drift, and output the final identity posterior probability.

[0105] In one possible implementation, the processing module 202 is used to divide the cross-device observation feature stream into sliding time windows under a unified clock and a unified coordinate reference, generate skeleton key point trajectories and optical flow directions from visual features, generate inertial measurement step frequency, ultra-wideband relative displacement increment and wireless channel state change from signal features, and fuse them to obtain velocity vector and unitized motion direction vector.

[0106] The processing module 202 takes the velocity vector and room-level topology map as input and combines them with the functional data of different candidate rooms to generate a accessibility score.

[0107] Processing module 202 is used to generate behavioral priors with time decay based on the room history entry records with a unified clock index and the current time period label as input.

[0108] The processing module 202 is used to calculate the room arrival probability distribution of each candidate room by using the normalized motion direction vector, accessibility score and behavior prior, and to calculate the estimated arrival time by combining the velocity vector and the path length of the room-level topology map.

[0109] In one possible implementation, the output module 203 is used to write the identity posterior probability and the room arrival probability into the decision context under a unified clock, and generate the target room index and the corresponding arrival time window.

[0110] The processing module 202 is used to decompose the device preparation delay by taking the target room index as input. The device preparation delay includes device discovery and authentication delay, media session pre-establishment delay, start-up buffer filling delay, and timeline alignment delay.

[0111] The processing module 202 is used to generate an active transfer command after comparing the device preparation delay with the expected arrival time. When the expected arrival time is less than the device preparation delay and the identity posterior probability and room arrival probability meet the trigger threshold, the processing module 202 generates an active transfer command.

[0112] In one possible implementation, the processing module 202 is used to trigger boundary event detection of the door frame area according to the active flow command. The boundary event detection generates arrival boundary markers and arrival timestamps based on infrared door magnetic triggers, ultra-wideband near-field threshold events and camera door frame intrusion events, and writes the arrival boundary markers and arrival timestamps into the decision context.

[0113] The processing module 202 is used to send mute and frame freeze commands through the distributed soft bus when the target room device receives the active streaming command and holds the media session identifier, and at the same time send unmute and playback unfreeze commands to the target room device.

[0114] The processing module 202 is used to control the current room device to stop audio output and freeze the rendering frame while the target room device completes playback unfreezing according to the timeline identifier.

[0115] In one possible implementation, the processing module 202 is used to complete device discovery, authentication and unified clock synchronization of the distributed device group through the distributed soft bus, and to establish a distributed clock and unified coordinate reference.

[0116] The processing module 202 is used to acquire human image video streams from the camera of the master device in the distributed device group and extract video frames through a key frame extraction algorithm.

[0117] The processing module 202 is used to input video frames into a deep learning model to generate user portrait re-identification embedding, skeleton key point motion vector and gait embedding, forming visual feature data.

[0118] The processing module 202 is used to collect wireless channel status information, Bluetooth signal strength trajectory, ultra-wideband ranging data and inertial measurement data from other devices in the distributed device group other than the master device, and form a signal characteristic data stream.

[0119] The processing module 202 is used to timestamp-align visual feature data and signal feature data using a distributed clock, and to normalize spatial location data based on a unified coordinate reference, and output a cross-device observation feature stream.

[0120] It should be noted that the above embodiments of the apparatus are only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0121] This embodiment also discloses an electronic device, as shown in the reference. Figure 3The electronic device may include: at least one processor 301, at least one communication bus 302, user interface 303, network interface 304, and at least one memory 305.

[0122] The communication bus 302 is used to enable communication between these components.

[0123] The user interface 303 may include a display screen and a camera. Optionally, the user interface 303 may also include a standard wired interface and a wireless interface.

[0124] The network interface 304 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0125] The processor 301 may include one or more processing cores. The processor 301 connects to various parts of the server using various interfaces and lines, and performs various server functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 305, and by calling data stored in memory 305. Optionally, the processor 301 may be implemented using at least one hardware form of Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), or Programmable Logic Array (PLA). The processor 301 may integrate one or a combination of several of the following: Central Processing Unit (CPU), Graphics Processing Unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications. The GPU is responsible for rendering and drawing the content required for display. The modem handles wireless communication. It is understood that the modem may also not be integrated into the processor 301 and may be implemented as a separate chip.

[0126] The memory 305 may include random access memory (RAM) or read-only memory. Optionally, the memory may include a non-transitory computer-readable storage medium. The memory 305 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), and instructions for implementing the various method embodiments described above. The data storage area may store data involved in the various method embodiments described above. Optionally, the memory 305 may also be at least one storage device located remotely from the aforementioned processor 301. As a computer storage medium, the memory 305 may include an operating system, a network communication module, a user interface 303 module, and an application program based on a facial recognition dynamic recommendation method.

[0127] exist Figure 3 In the illustrated electronic device, the user interface 303 is primarily used to provide an input interface for the user and to acquire user input data. The processor 301 can be used to call an application program stored in the memory 305 that uses a dynamic recommendation method based on facial recognition. When executed by one or more processors 301, the electronic device performs one or more methods as described in the above embodiments.

[0128] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the present invention is not limited to the described order of actions, as some steps can be performed in other orders or simultaneously according to the present invention. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to the present invention.

[0129] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0130] In the several embodiments provided by this invention, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some service interface; the indirect coupling or communication connection between apparatuses or units may be electrical or other forms.

[0131] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0132] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0133] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory 305 and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned memory 305 includes various media capable of storing program code, such as a USB flash drive, external hard drive, magnetic disk, or optical disk.

[0134] The present invention also discloses a computer-readable storage medium storing instructions. When executed by one or more processors 301, these instructions cause an electronic device to perform one or more methods as described in the above embodiments.

[0135] The above are merely exemplary embodiments of this disclosure and should not be construed as limiting the scope of this disclosure. Any equivalent changes and modifications made in accordance with the teachings of this disclosure shall still fall within the scope of this disclosure. Those skilled in the art will readily conceive of other embodiments of this disclosure upon considering the specification and the disclosure of practical truths. This invention is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not described in this disclosure. The specification and embodiments are to be considered exemplary only, and the scope and spirit of this disclosure are defined by the claims.

Claims

1. A dynamic recommendation method based on facial recognition, characterized in that, The method includes: The user's visual feature data is obtained by processing the user's portrait video stream and aligning it with the signal feature data shared by the distributed device group to output a cross-device observation feature stream; Based on the visual feature data and combined with the user's historical preference data, target recommended content is generated; Using the cross-device observation feature stream as input, the posterior probability of the user's identity is calculated based on log-likelihood ratio weighted fusion; Based on the posterior probability of identity and the cross-device observation feature stream, the user's migration intention is calculated based on the direction of movement, accessibility, and behavioral prior, and the room arrival probability distribution and estimated arrival time of different candidate rooms are output. When both the identity posterior probability and the room arrival probability of the target room meet the trigger threshold, and the estimated arrival time of the target room is less than the device preparation delay of the target room, an active transfer command is output. According to the active flow command, when the user is detected to have reached the boundary of the target room, the target recommended content is flowed to the device in the target room.

2. The dynamic recommendation method based on facial recognition according to claim 1, characterized in that, The process of generating target recommended content based on the visual feature data and combined with user historical preference data specifically includes: After the visual feature data is standardized and mapped and aligned, it is spatially aligned with the historical embedding vector corresponding to the user's historical preference data to obtain the visual embedding and the historical embedding. The visual embedding and the historical embedding are dynamically adjusted based on the intensity and consistency of the visual features to determine the weights of short-term interests and long-term preferences, and a personalized state vector is output. By combining the personalized state vector with the age classification information and scene context information determined based on the visual feature data, a constraint mask is generated to filter out the content and obtain candidate content. The personalized state vector is matched and scored with the candidate content vector, and the original score set of candidate content is generated by combining the session context vector and the platform quality factor. Based on the original score set, probability normalization is performed under the constraint mask to obtain target recommended content that meets the display conditions.

3. The dynamic recommendation method based on facial recognition according to claim 1, characterized in that, The step of calculating the user's posterior probability of identity based on log-likelihood ratio weighted fusion, using the cross-device observation feature stream as input, specifically includes: The cross-device observation feature stream is processed for time and space alignment using a unified clock and a unified coordinate reference, and the output is a feature tensor sequence organized by time window. For the visual features and signal features in the feature tensor sequence, feature reliability weights are generated respectively, and the feature reliability weights are adaptively updated based on consistency score, drift score and signal-to-noise ratio score; Based on the feature tensor sequence and the feature reliability weights, modal likelihood models are established under user conditions and non-user conditions, respectively, and the modal level log-likelihood ratio sequence is output. The modal-level log-likelihood ratio sequence and the feature reliability weights are weighted and accumulated to form an immediate posterior probability of identity. The real-time posterior probability of identity and the posterior probability of identity in the previous time window are subjected to temporal smoothing, and the smoothed posterior probability of identity and the updated temporal fusion coefficient are output. When the smoothed posterior identity probability is lower than the confidence threshold, the modality selection and re-estimation process is triggered to suppress modalities with significant drift and output the final posterior identity probability.

4. The dynamic recommendation method based on facial recognition according to claim 1, characterized in that, The process of calculating the user's migration intention based on the posterior probability of identity and the cross-device observation feature stream, using motion direction, accessibility, and behavioral priors, and outputting the room arrival probability distribution and estimated arrival time for different rooms, specifically includes: The cross-device observation feature stream is divided into sliding time windows under a unified clock and a unified coordinate reference. The skeleton key point trajectory and optical flow direction are generated from the visual features. The inertial measurement step frequency, ultra-wideband relative displacement increment and wireless channel state change are generated from the signal features. The velocity vector and the unitized motion direction vector are fused together. The velocity vector and room-level topology map are used as inputs, and combined with the functional data of different candidate rooms, a accessibility score is generated. Based on the room history entry records and the current time period label of the unified clock index as input, a behavioral prior with time decay is generated. The room arrival probability distribution of each candidate room is calculated using the unitized motion direction vector, the accessibility score, and the behavioral prior, and the estimated arrival time is calculated by combining the velocity vector and the path length of the room-level topology map.

5. The dynamic recommendation method based on facial recognition according to claim 1, characterized in that, When both the identity posterior probability and the room arrival probability of the target room meet the trigger threshold, and the estimated arrival time of the target room is less than the device preparation delay of the target room, an active transfer command is output, specifically including: The posterior probability of the identity and the probability of arriving at the room are written into the decision context under a unified clock, and the target room index and the corresponding arrival time window are generated. The target room index is invoked as input to decompose the device preparation delay, which includes device discovery and authentication delay, media session pre-establishment delay, start-up buffer filling delay, and timeline alignment delay. After comparing the device preparation delay with the estimated arrival time, if the estimated arrival time is less than the device preparation delay and the identity posterior probability and the room arrival probability meet the trigger threshold, an active transfer command is generated.

6. The dynamic recommendation method based on facial recognition according to claim 1, characterized in that, The device that, based on the active transfer command, transfers the target recommended content to the target room when it detects that the user has reached the boundary of the target room, specifically includes: The active flow command triggers boundary event detection in the door frame area. The boundary event detection generates arrival boundary markers and arrival timestamps based on infrared door magnetic triggers, ultra-wideband near-field threshold events and camera door frame intrusion events, and writes the arrival boundary markers and arrival timestamps into the decision context. When the target room device receives the active streaming command and holds a media session identifier, it sends mute and frame freeze commands through the distributed soft bus, and simultaneously sends unmute and playback unfreeze commands to the target room device. While the target room device completes playback unfreezing according to the timeline markers, the current room device is controlled to stop audio output and freeze the rendering frame.

7. The dynamic recommendation method based on facial recognition according to claim 1, characterized in that, The process of acquiring and processing the user's facial video stream to obtain the user's visual feature data, aligning it with the signal feature data shared by the distributed device group, and outputting a cross-device observation feature stream specifically includes: The distributed soft bus is used to complete the device discovery, authentication and unified clock synchronization of the distributed device group, and to establish a distributed clock and unified coordinate reference. The human image video stream is captured by the camera of the master device in the distributed device group, and video frames are extracted by a key frame extraction algorithm. The video frames are input into a deep learning model to generate the user's portrait re-identification embedding, skeleton key point motion vectors and gait embedding, forming the visual feature data; The wireless channel status information, Bluetooth signal strength trajectory, ultra-wideband ranging data and inertial measurement data are collected by other devices in the distributed device group other than the main device to form the signal feature data stream; The visual feature data and the signal feature data are timestamped using the distributed clock, and the spatial location data is normalized based on the unified coordinate reference to output the cross-device observation feature stream.

8. A dynamic recommendation device based on facial recognition, characterized in that, The device is used to execute a dynamic recommendation method based on facial recognition as described in any one of claims 1-7, and the device includes an acquisition module, a processing module, and an output module, wherein: The acquisition module is used to process the user's visual feature data obtained by collecting the user's portrait video stream, align it with the signal feature data shared by the distributed device group, and output a cross-device observation feature stream. The processing module is used to generate target recommended content based on the visual feature data and combined with the user's historical preference data; The processing module is used to calculate the posterior probability of the user's identity based on the weighted fusion of the cross-device observation feature stream as input. The processing module is used to calculate the user's migration intention based on the identity posterior probability and the cross-device observation feature stream, and based on the direction of movement, accessibility, and behavioral priors, and output the room arrival probability distribution and estimated arrival time for different candidate rooms; The processing module is configured to output an active transfer command when both the identity posterior probability and the room arrival probability of the target room meet the trigger threshold, and the estimated arrival time of the target room is less than the device preparation delay of the target room. The output module is used to transfer the target recommended content to the device in the target room when the user is detected to have reached the boundary of the target room, according to the active transfer command.

9. An electronic device, characterized in that, The device includes a processor, a communication bus, a user interface, a network interface, and a memory. The memory is used to store instructions. The user interface and the network interface are both used to communicate with other devices. The communication bus is used to enable communication between the components within the electronic device. The processor is used to execute the instructions stored in the memory to cause the electronic device to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed, perform the method as described in any one of claims 1-7.