Cross-screen visual content generation method and apparatus, device, storage medium and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CHINA R&D CENT
- Filing Date
- 2026-05-19
- Publication Date
- 2026-08-07
AI Technical Summary
但是这种方法不仅耗时耗力,更严重制约了视觉内容大规模部署的实施效率
[0009]本公开实施例的跨屏视觉内容生成方法和装置、设备、存储介质和程序产品,通过对原始视觉内容进行语义分析,确定原始视觉内容中视觉对象的原始布局及各视觉对象的权重值;根据目标显示设备的参数和各视觉对象的权重值调整原始布局,确定与原始视觉内容的整体语义特征一致的目标布局;根据目标布局,生成在目标显示设备上显示的包含原始视觉内容中视觉对象的新视觉内容;利用对原始视觉内容的语义分析,可以深度解析原始视觉内容的语义内涵,从而理解原始视觉内容的意图,可以生成与原始视觉内容的意图最接近新视觉内容,可以实现根据目标显示设备对视觉内容的自适应调整和优化呈现,可以在保持与原始视觉内容的意图一致性的基础上,实现视觉内容跨多终端屏幕的生成与分发,可以使视觉内容具备“一次创作,多端适配”的批量发布能力,可以提高对视觉内容大规模部署的实施效率。特别是对于挖孔屏幕和可变尺寸屏幕等异形屏幕,可以在确保原始视觉内容的意图不发生改变的基础上,对原始视觉内容的布局进行动态调整,可以保证原始视觉内容的关键信息不丢失,可以适应屏幕尺寸的变化,在变化的屏幕尺寸中都能够获得最佳的呈现效果,从而可以提升用户的视觉体验。
Smart Images

Figure CN122529959A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method for generating cross-screen visual content, a device for generating cross-screen visual content, an electronic device, a computer storage medium, and a computer program product. Background Technology
[0002] Currently, cross-screen display of visual content is prone to problems such as display distortion or black borders due to differences in screen size. For punch-hole screens, key information in the visual content may be located precisely within the punch-hole area, resulting in the inability to display this information correctly. For variable-size screens, the inability to dynamically adjust the visual content according to the screen size in real time leads to suboptimal presentation. To overcome these problems, multiple sets of visual content are typically created and adapted for each screen. However, this approach is not only time-consuming and labor-intensive but also severely restricts the efficiency of large-scale deployment of visual content. Summary of the Invention
[0003] This disclosure provides a method, apparatus, device, storage medium, and program product for generating cross-screen visual content.
[0004] According to a first aspect, embodiments of this disclosure provide a cross-screen visual content generation method, comprising: determining the original layout of visual objects in the original visual content and the weight values of each visual object based on semantic analysis of the original visual content; adjusting the original layout based on parameters of a target display device and the weight values to determine a target layout consistent with the overall semantic features of the original visual content; and generating new visual content containing the visual objects in the original visual content for display on the target display device based on the target layout.
[0005] According to a second aspect, embodiments of this disclosure provide a cross-screen visual content generation apparatus, comprising: a semantic analysis module configured to perform semantic analysis on original visual content to determine the original layout of visual objects in the original visual content and the weight values of each visual object; a layout adjustment module configured to adjust the original layout based on parameters of a target display device and the weight values to determine a target layout consistent with the overall semantic features of the original visual content; and a content generation module configured to generate new visual content containing the visual objects in the original visual content for display on the target display device based on the target layout.
[0006] According to a third aspect, embodiments of this disclosure provide an electronic device including one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the methods described in the first aspect.
[0007] According to a fourth aspect, embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon that, when executed by a processor, implements the methods described in the first aspect.
[0008] According to a fifth aspect, embodiments of this disclosure provide a computer program product including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0009] The cross-screen visual content generation method, apparatus, device, storage medium, and program product of this disclosure perform semantic analysis on the original visual content to determine the original layout of visual objects and the weight values of each visual object in the original visual content; adjust the original layout according to the parameters of the target display device and the weight values of each visual object to determine a target layout consistent with the overall semantic features of the original visual content; and generate new visual content containing the visual objects of the original visual content for display on the target display device based on the target layout. By utilizing the semantic analysis of the original visual content, the semantic connotation of the original visual content can be deeply analyzed, thereby understanding the intent of the original visual content. This allows for the generation of new visual content that is closest to the intent of the original visual content, enabling adaptive adjustment and optimized presentation of the visual content according to the target display device, while maintaining consistency with the intent of the original visual content. Figure 1 Building upon consistency, enabling the generation and distribution of visual content across multiple terminal screens allows for the batch release of visual content with the capability of "creating once and adapting to multiple platforms," thereby improving the efficiency of large-scale deployment. Particularly for irregularly shaped screens such as punch-hole displays and variable-size screens, the layout of the original visual content can be dynamically adjusted while ensuring the intent remains unchanged. This guarantees that key information is not lost and adapts to changes in screen size, achieving optimal presentation across varying screen dimensions, thus enhancing the user's visual experience.
[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0011] Figure 1 This is an exemplary system architecture diagram of an embodiment of the cross-screen visual content generation method disclosed herein; Figure 2 This is a flowchart of a cross-screen visual content generation method according to some embodiments of this disclosure; Figure 3A This is a system architecture diagram of one embodiment of the cross-screen visual content generation method disclosed herein; Figure 3B This is a schematic diagram illustrating the generation of new visual content based on a target layout, according to another embodiment of this disclosure; Figure 4 This is a flowchart of some embodiments of the present disclosure performing semantic analysis on the original visual content; Figure 5 This is a flowchart illustrating some embodiments of the present disclosure that determine a target layout based on device parameters and weight values; Figure 6A This is a schematic diagram illustrating the generation of candidate layout groups through a layout generation strategy according to an embodiment of this disclosure; Figure 6B This is a schematic diagram illustrating the iterative optimization of the layout generation strategy according to an embodiment of this disclosure; Figure 7 This is a flowchart of a cross-screen visual content generation method according to some other embodiments of this disclosure; Figure 8A This is a schematic diagram illustrating the weighted sorting of multidimensional weight values according to an embodiment of this disclosure; Figure 8B This is a schematic diagram illustrating layout adjustment triggered by edge computing, representing an embodiment of this disclosure. Figure 9 This is a schematic diagram illustrating the repair of occlusion and expansion according to an embodiment of this disclosure; Figure 10 This is a schematic diagram of the first application scenario using the cross-screen visual content generation method disclosed herein; Figure 11 This is a schematic diagram of a second application scenario using the cross-screen visual content generation method disclosed herein; Figure 12 This is a schematic diagram of a third application scenario using the cross-screen visual content generation method disclosed herein; Figure 13 This is a schematic diagram of the fourth application scenario using the cross-screen visual content generation method disclosed herein; Figure 14 This is a schematic diagram of one embodiment of the cross-screen visual content generation apparatus according to the present disclosure; Figure 15 This is a schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure. Detailed Implementation
[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0013] It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other. This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0014] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0015] Figure 1 An exemplary system architecture 100 is shown, to which embodiments of the cross-screen visual content generation method of this disclosure can be applied.
[0016] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, network 104, and server 105. Network 104 serves as a medium for providing communication links between terminal devices 101, 102, and 103 and server 105, and between the terminal devices themselves. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables.
[0017] Users can use terminal devices 101, 102, and 103 to interact with other terminal devices or servers 105 via network 104 to receive or send messages, etc. Client application software, such as visual content playback applications and communication applications, can be installed on terminal devices 101, 102, and 103.
[0018] Terminal devices 101, 102, and 103 can be either hardware or software. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices, including but not limited to televisions, laptops, tablets, mobile phones, and projection devices. When terminal devices 101, 102, and 103 are software, they can be installed in the electronic devices listed above. They can be implemented as multiple software programs or software modules, or as a single software program or software module. No specific limitations are made here.
[0019] Server 105 can be a server providing various services, such as a visual content management server providing visual content management services. Server 105 can send the same visual content to terminal devices 101, 102, and 103 respectively, and play it on terminal devices 101, 102, and 103 respectively. Among them, terminal devices 101, 102, and 103 can be target display devices. Before sending the visual content to terminal devices 101, 102, and 103, server 105 can also generate original visual content and perform the following operations based on the original visual content: based on semantic analysis of the original visual content, determine the original layout of visual objects in the original visual content and the weight value of each visual object; adjust the original layout based on the parameters and weight values of the target display device to determine a target layout consistent with the overall semantic features of the original visual content; based on the target layout, generate new visual content containing the visual objects in the original visual content that is displayed on the target display device.
[0020] It should be noted that server 105 can be either hardware or software. When server 105 is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When server 105 is software, it can be implemented as multiple software programs or software modules (e.g., used to provide visual content generation and management services), or as a single software program or software module. No specific limitations are made here.
[0021] It should be noted that the cross-screen visual content generation method provided in the embodiments of this disclosure can be executed by server 105, or by server 105 and terminal devices 101, 102, and 103 in cooperation with each other. Accordingly, the various parts (e.g., modules and sub-modules) included in the cross-screen visual content generation device can all be set in server 105, or they can be set in server 105 and terminal devices 101, 102, and 103 respectively.
[0022] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0023] Figure 2 A flow 200 of a cross-screen visual content generation method according to some embodiments of this disclosure is shown. For example... Figure 2 As shown, process 200 may include the following steps: Step 201: Based on semantic analysis of the original visual content, determine the original layout of visual objects and the weight value of each visual object in the original visual content.
[0024] In this embodiment, the execution entity (e.g., Figure 1The server in the process can perform semantic analysis on the original visual content to determine the original layout of visual objects and the weight values of each visual object. Visual content can refer to various information and forms of expression conveyed visually. For example, visual content can include images, videos, animations, etc., and the embodiments of this disclosure do not limit the type of visual content. The original visual content can be visual content generated by the executing entity, visual content obtained by the executing entity from locally stored visual content, or visual content obtained by the executing entity from other devices via a network. The embodiments of this disclosure do not limit the method of obtaining the original visual content.
[0025] In this embodiment, the original visual content may include visual objects, which can refer to distinguishable entities within the original visual content that have independent semantic meaning. For example, such as Figure 3A As shown, the original visual content 301 is an advertising poster for a mobile phone, which includes a mobile phone trademark 302, mobile phone selling points 303, front of the mobile phone 304, and back of the mobile phone 305. Among them, the mobile phone trademark 302, mobile phone selling points 303, front of the mobile phone 304, and back of the mobile phone 305 are the visual objects in the original visual content 301.
[0026] In this embodiment, the executing entity performs semantic analysis on the original visual content to deeply analyze its semantic connotations and understand its intent. The intent of the original visual content may include, but is not limited to, the purpose of the image, the creator's creative concept, etc., and can be presented through the original layout and weight values of visual objects within the original visual content. The original layout of visual objects refers to the way they are organized, arranged, and configured within the original visual content. The weight values of visual objects refer to their importance within the overall semantics of the original visual content.
[0027] In this embodiment, the executing entity can extract features from the original visual content, perform multi-level semantic analysis on the original visual content, determine the original layout of visual objects and the weight values of each visual object, and thus understand the intent of the original visual content. For example, multi-level semantic analysis can include object-level semantic analysis, relation-level semantic analysis, and scene-level semantic analysis, which can obtain the meaning of each visual object in the original visual content, the meaning of the relationships between visual objects, and the meaning of the entire visual content scene.
[0028] In an optional example, such as Figure 3AAs shown, the original visual content 301 is input into the semantic analysis module 306, which performs feature extraction and semantic understanding on the original visual content 301, realizing multi-level semantic analysis of the original visual content 301. The multi-level semantic analysis performed by the semantic analysis module 306 on the original visual content 301 may include: object-level detection and segmentation, scene-level analysis and understanding, and determination of relation-level regions of interest (ROIs). The semantic analysis module 306 can output the original layout of visual objects in the original visual content 301 and the weight values 307 of each visual object.
[0029] Step 202: Adjust the original layout based on the parameters and weight values of the target display device to determine a target layout that is consistent with the overall semantic features of the original visual content.
[0030] In this embodiment, the aforementioned execution entity can adjust the original layout of visual objects according to the parameters of the target display device and the weight values of each visual object, thereby determining a target layout consistent with the overall semantic features of the original visual content. The target display device can be a device used to display the original visual content; it can be a flat display device or a stereoscopic display device; it can be a device with a standard screen or a device with a non-standard screen. For example, a non-standard screen can include a punch-hole screen, a variable-size screen, or other irregularly shaped screens. The embodiments of this disclosure do not limit the type of target display device. The parameters of the target display device can include display size, resolution, irregularly shaped screen space matrix, etc., and these parameters can be determined according to the type of target display device. The embodiments of this disclosure do not limit this.
[0031] In this embodiment, the overall semantic feature refers to a high-level semantic representation abstracted from the entire visual content. It can reflect the object categories, global scene meaning, entity relationships, behavioral events, and emotional style of the visual content, approximating the overall human understanding of visual content. The executing entity can adjust the original layout of visual objects in the original visual content based on the parameters of the target display device and the weight values of the visual objects in the original visual content to obtain a target layout that meets the display requirements of the target display device. The overall semantic feature of the target layout is consistent with the overall semantic feature of the original visual content; consistency here can also include cases where the overall semantic features are similar.
[0032] In this embodiment, semantic analysis is performed on the adjusted layout to determine its overall semantic features. These features are then compared with the overall semantic features of the original visual content. The adjusted layout that best matches the overall semantic features of the original visual content is selected as the target layout. The adjustment to the original layout causes a change in the weight values of the visual objects. The overall semantic features of the adjusted layout include a comprehensive abstraction of both the adjusted layout and the changed weight values.
[0033] In an optional example, such as Figure 3A As shown, the parameters 308 of the target display device, the original layout of the visual object, and the weight values 307 of each visual object are input into the layout adjustment module 309 to adjust the original layout of the visual object. A target layout 310, consistent with the overall semantic features of the original visual content 301, is then determined as the output of the layout adjustment module 309. During the adjustment of the original layout, the layout adjustment module 309 can intelligently cut and expand components according to the parameters 308 of the target display device to adapt to its shape.
[0034] When the target display device's screen size is larger than the original visual content size, such as Figure 3A As shown, when the screen width of the target display device is greater than the width of the original visual content, the layout adjustment module 309 can expand components during the adjustment of the original layout, so that the resulting target layout 310 is adapted to the screen width of the target display device. When the screen size of the target display device is smaller than the size of the original visual content, such as... Figure 3B As shown, the screen width of the target display device is smaller than the width of the original visual content. During the process of adjusting the original layout, the layout adjustment module 309 can perform intelligent cutting so that the target layout 311 is adapted to the screen width of the target display device.
[0035] Step 203: Based on the target layout, generate new visual content that includes visual objects from the original visual content and is displayed on the target display device.
[0036] In this embodiment, the execution entity can generate new visual content, containing visual objects from the original visual content, for display on the target display device, based on the target layout. Both the target layout and the original layout can be mask images. The execution entity can generate new visual content, containing visual objects from the original visual content, for display on the target display device based on the mask image of the target layout and the original visual content. The generation of new visual content from the target layout can be achieved using a generative artificial intelligence model, such as a diffusion model. The embodiments of this disclosure do not limit the implementation method of generating new visual content based on the target layout.
[0037] In an optional example, such as Figure 3A As shown, the target layout 310 obtained by the layout adjustment module 309 is input into the content generation module 312. The content generation module 312 can generate new visual content 313 containing visual objects from the original visual content for display on the target display device, based on the target layout 310 and the original visual content 301. During the generation of the new visual content 313, the content generation module 312 can also perform occlusion and / or expansion repair on the new visual content 313 based on the original visual content 301, and handle style consistency processing, so that the resulting new visual content 313 can maintain semantic integrity and style consistency with the original visual content 301.
[0038] In another optional example, such as Figure 3B As shown, the prompt word 315, the original visual content 301, and the target layout 311 obtained by the layout adjustment module 309 are input into the content generation module 312. Under the guidance of the prompt word 315, the content generation module 312 can generate new visual content 314 containing visual objects in the original visual content for display on the target display device, based on the target layout 311 and the original visual content 301.
[0039] Both new visual content 313 and 314 are generated based on original visual content 301 and contain the same visual objects as original visual content 301. Original visual content 301 is an advertising poster designed according to the parameters of a standard screen. New visual content 313 can be visual content generated from original visual content 301 that conforms to the parameters of another standard screen, or it can be visual content generated from original visual content 301 that conforms to the parameters of a non-standard screen. The display device on which new visual content 313 is applied has a different screen ratio than the display device on which original visual content 301 is applied. New visual content 314 can be visual content generated from original visual content 301 that conforms to the parameters of a non-standard irregularly shaped screen. The display device on which new visual content 314 is applied has a circular screen, while the display device on which original visual content 301 is applied has a rectangular screen.
[0040] The cross-screen visual content generation method provided in this disclosure performs semantic analysis on the original visual content to determine the original layout of visual objects and the weight values of each visual object in the original visual content; adjusts the original layout according to the parameters of the target display device and the weight values of each visual object to determine a target layout consistent with the overall semantic features of the original visual content; and generates new visual content containing the visual objects of the original visual content for display on the target display device based on the target layout. By utilizing semantic analysis of the original visual content, the semantic connotation of the original visual content can be deeply analyzed, thereby understanding the intent of the original visual content. This allows for the generation of new visual content that is closest to the intent of the original visual content, enabling adaptive adjustment and optimized presentation of the visual content according to the target display device, while maintaining consistency with the intent of the original visual content. Figure 1 Building upon consistency, enabling the generation and distribution of visual content across multiple terminal screens allows for the batch release of visual content with the capability of "creating once and adapting to multiple platforms," thereby improving the efficiency of large-scale deployment. Particularly for irregularly shaped screens such as punch-hole displays and variable-size screens, the layout of the original visual content can be dynamically adjusted while ensuring the intent remains unchanged. This guarantees that key information is not lost and adapts to changes in screen size, achieving optimal presentation across varying screen dimensions, thus enhancing the user's visual experience.
[0041] Figure 4 The flowchart illustrating some embodiments of this disclosure demonstrates the process of performing semantic analysis on raw visual content. For example... Figure 4 As shown, determining the original layout of visual objects and the weight values of each visual object based on semantic analysis of the original visual content can include the following steps: Step 401: Perform instance segmentation on the original visual content to determine the category, bounding box, and mask of the visual objects in the original visual content.
[0042] In this embodiment, the execution entity (e.g., Figure 1 Server 105 can perform instance segmentation on the original visual content, determining the category, bounding box, and mask of visual objects within the original visual content. Instance segmentation is a computer vision task aimed at performing pixel-level precise segmentation of each object instance in an image and assigning a unique identifier to each individual instance. Instance segmentation of an image yields the category, bounding box, and mask of the objects within it. The object category can be the category label to which each segmented object belongs. The object's bounding box can be a rectangle surrounding each object, used to locate the object's position in the image. The object's mask can be the pixel-level segmentation result, used to accurately depict the object's outline and shape.
[0043] In this embodiment, the executing entity can input the original visual content into an artificial intelligence model, extract features from the original visual content, and perform instance segmentation within the original visual content based on the extracted features, outputting the category, bounding box, and mask of the visual objects in the original visual content. For example, the artificial intelligence model used for instance segmentation can use an improved MaskDINO model, which can be a deep learning model combining DINO self-supervised learning and Mask R-CNN. The embodiments of this disclosure do not limit the implementation method for instance segmenting the original visual content.
[0044] In an optional example, the execution entity can input the raw visual content into the improved MaskDINO model, perform pixel-level instance segmentation on the raw visual content, and output structured data containing the categories, bounding boxes, and masks of the visual objects in the raw visual content. The improved MaskDINO model can also visualize the structured data, outputting a mask image of the visual objects in the raw visual content, as well as an image annotating the categories and bounding boxes of the visual objects in the raw visual content.
[0045] Step 402: Based on the category, bounding box, and mask of the visual object, perform scene semantic analysis on the original visual content to determine the scene semantic features.
[0046] In this embodiment, the aforementioned execution entity can perform scene semantic analysis on the original visual content based on the category, bounding box, and mask of the visual objects to determine the scene semantic features. Scene semantic analysis, based on the identified basic objects, further understands the spatial, temporal, and logical relationships between them and integrates multi-source contextual information to achieve a high-level semantic interpretation of the entire scene. The execution entity can input the category, bounding box, and mask of the visual objects obtained from the original visual content into the artificial intelligence model, and perform scene semantic analysis on the original visual content using the category, bounding box, and mask, outputting the scene semantic features of the original visual content.
[0047] In some optional implementations, step 402, based on the category, bounding box, and mask of the visual object, performs scene semantic analysis on the original visual content to determine scene semantic features. This may include: extracting image and text features from the category, bounding box, and mask of the visual object, performing feature alignment, calculating the similarity after feature alignment, and determining the scene label of the original visual content; and performing scene semantic reasoning on the scene label based on a knowledge graph to determine the scene semantic features of the original visual content. The extraction, alignment, and similarity calculation of image and text features can be performed using a cross-modal model. For example, the CLIP model can be used as a cross-modal model. The implementations of this disclosure do not limit the type of cross-modal model.
[0048] In an optional example, the agent can input the visualization output of the visual object's category, bounding box, and mask obtained from the original visual content into the CLIP model to extract image and text features. The image and text features are then aligned, and the cosine similarity between the aligned features is calculated. The feature with the highest similarity is selected as the scene label. The agent can then perform scene semantic reasoning based on the scene label and a knowledge graph to obtain the scene semantic features of the original visual content. The knowledge graph can be ConceptNet.
[0049] Step 403: Based on the category, bounding box, and mask of the visual object and the semantic features of the scene, determine the original layout of the visual object and the weight value of each visual object.
[0050] In this embodiment, the aforementioned execution entity can determine the original layout of visual objects and the weight values of each visual object based on the category, bounding box, and mask of the visual objects and scene semantic features. Specifically, the original layout of visual objects in the original visual content can be determined based on their category, bounding box, and mask. A relationship graph of visual objects in the original visual content can be constructed based on their category, bounding box, and mask, and scene semantic features, and the weight values of each visual object in the original visual content can be determined based on this relationship graph.
[0051] In some optional implementations, step 403, based on the category, bounding box, and mask of the visual object and scene semantic features, determines the original layout of the visual object and the weight value of each visual object. This may include: determining the original layout of the visual object based on its category, bounding box, and mask; constructing a visual object relationship graph based on the category, bounding box, and mask of the visual object and scene semantic features; and determining the weight value of each visual object based on the visual object relationship graph. For example, a graph attention network (GAT) can be used to construct the visual object relationship graph, and the implementations of this disclosure are not limited thereto.
[0052] In an optional example, the implementing entity can use GAT to construct a relationship graph of visual objects in the original visual content. First, nodes in the relationship graph can be determined based on the categories of visual objects. Then, connections between nodes are constructed based on the categories, bounding boxes, masks, and scene semantic features of the visual objects, forming a graph structure. Next, the correlation between different visual objects is learned through GAT's attention mechanism, obtaining the attention weights of each node in the graph structure relative to other nodes. These weights reflect the relative importance of the node in the graph. These attention weights can then be mapped onto the corresponding visual objects, forming an attention heatmap. Based on the attention heatmap, the region of interest for each visual object and the weight values of each visual object can be determined.
[0053] This embodiment performs multi-level semantic analysis on the original visual content, which can accurately identify the composition of the original visual content, deeply analyze the semantic connotation of the original visual content, and provide data support for adjusting the layout according to the parameters of the target display device.
[0054] Figure 5 The following illustrates a process for determining a target layout based on device parameters and weight values, according to some embodiments of this disclosure. For example... Figure 5 As shown, adjusting the original layout based on the parameters and weight values of the target display device to determine a target layout consistent with the overall semantic features of the original visual content may include the following steps: Step 501: Using a layout generation strategy, adjust the original layout according to the parameters and weight values of the target display device to generate a candidate layout group containing multiple candidate layouts.
[0055] In this embodiment, the execution entity (e.g., Figure 1 Server 105 can use a layout generation strategy to adjust the original layout based on the parameters and weight values of the target display device, generating a candidate layout group containing multiple candidate layouts. The layout generation strategy can be a pre-defined strategy that adjusts the original layout of visual objects based on the parameters of the target display device and the weight values of the visual objects. For example, the layout generation strategy may include determining how to adjust the layout based on the priority of the weights; adopting a balanced display mode to ensure a harmonious visual presentation of the overall layout; performing weight balancing calculations for multiple object regions to ensure that the weights of all objects meet priority requirements while maintaining the rationality of the overall layout; and combining this with: scaling core objects proportionally, appropriately scaling and cropping secondary objects, and filling missing areas with layout fill. The core and secondary objects are determined based on the priority of the visual object weights, and the priority of the weights is determined by the magnitude of the weight values.
[0056] In this embodiment, by adjusting the original layout using a layout generation strategy, a candidate layout group comprising multiple candidate layouts can be generated. Each candidate layout includes the same visual objects as the original layout, but has a different organization, arrangement, and configuration of these visual objects. The weight values of each visual object in each candidate layout may differ from those in the original layout, and the visual objects are ordered according to their weight values, and may be the same as or similar to those in the original layout. The organization, arrangement, and configuration of visual objects differ among the candidate layouts in the candidate layout group, and the weight values of the same visual objects within each candidate layout may differ.
[0057] In an optional example, such as Figure 6A As shown, the original layout 601 includes four semantic objects: a first semantic object 602, a second semantic object 603, a third semantic object 604, and a fourth semantic object 605. The weight values of the four semantic objects decrease sequentially. A semantic object library 606 can be constructed based on the semantic objects in the original layout 601 and their corresponding weight values. Semantic objects and their corresponding weight values can be obtained from the semantic object library 606 and combined with the new display parameters 607 of the target display device. The original layout 601 is then adjusted using a layout generation strategy 608 to generate a candidate layout group 609 containing multiple candidate layouts. The behavior 608c of adjusting visual objects can include moving, scaling, cropping, rotating, and flipping semantic objects.
[0058] Step 502: Based on semantic analysis of the candidate layouts and comparison with the overall semantic features of the original visual content, determine multiple target candidate layouts whose differences from the overall semantic features are less than a difference threshold, and determine the target layout from the multiple target candidate layouts.
[0059] In this embodiment, the execution entity can perform semantic analysis on the candidate layouts and compare them with the overall semantic features of the original visual content to determine multiple target candidate layouts whose differences from the overall semantic features are less than a preset difference threshold. The target layout is then determined from these multiple target candidate layouts. The target candidate layout can be a candidate layout consistent with the overall semantic features of the original visual content, or it can be randomly selected from multiple target candidate layouts as the target layout. In this embodiment, consistency between the target candidate layout and the overall semantic features of the original visual content can mean that the difference between the target candidate layout and the overall semantic features of the original visual content is less than a preset difference threshold.
[0060] In some alternative implementations, semantic analysis of candidate layouts within a candidate layout group, and comparison with the overall semantic features of the original visual content, can be performed using a system that performs semantic analysis on the original visual content. The candidate layout group can be used as feedback input to perform feedback analysis on the system that performs semantic analysis on the original visual content, thereby filtering out target candidate layouts that are consistent with the overall semantic features of the original visual content.
[0061] The system for semantic analysis of the original visual content can be a multimodal semantic analysis system that performs multi-level semantic analysis of the original visual content. Based on the semantic analysis of the original visual content, determining the original layout of visual objects and the weight values of each visual object in the original visual content can include: performing semantic analysis on the original visual content using the multimodal semantic analysis system to determine the original layout of visual objects and the weight values of each visual object. Step 502, based on semantic analysis of candidate layouts and comparison with the overall semantic features of the original visual content, determines multiple target candidate layouts whose differences from the overall semantic features are less than a difference threshold. This can include: inputting the candidate layout group into the multimodal semantic analysis system, performing semantic analysis on the candidate layouts, comparing them with the overall semantic features of the original visual content, and determining target candidate layouts whose differences from the overall semantic features are less than a difference threshold.
[0062] In some optional implementations, adjusting the original layout based on the parameters and semantic weight values of the target display device to determine a target layout consistent with the overall semantic features of the original visual content may further include: updating the layout generation strategy based on the target candidate layouts; and adjusting the original layout according to the parameters and weight values of the target display device using the updated layout generation strategy to generate a candidate layout group containing multiple candidate layouts. Subsequently, a step can be performed to perform semantic analysis on the candidate layouts and compare them with the overall semantic features of the original visual content to determine multiple target candidate layouts whose differences from the overall semantic features are less than a difference threshold. This forms an iterative "create-evaluate-optimize" process, continuously optimizing the layout generation strategy to generate target candidate layouts that are closer to the overall semantic features of the original visual content.
[0063] In an optional example, such as Figure 6BAs shown, the layout generation strategy can be a layout generation strategy model 610. A candidate layout group 611, containing multiple candidate layouts, generated according to the layout generation strategy model 610, can be input into a multimodal semantic analysis system 612. Semantic analysis is performed on each candidate layout in the candidate layout group 611, and compared with the overall semantic features of the original visual content. From this, a target candidate layout group that better matches the overall semantic features of the original visual content is selected. The target candidate layout group is then input back into the layout generation strategy model 610 for fine-tuning. Based on the fine-tuned layout generation strategy model 610, a candidate layout group 611 is generated, resulting in a layout-optimized candidate layout group. A new target candidate layout group 614 can be determined from the layout-optimized candidate layout group, and the target layout can be determined from the new target candidate layout group 614. The element integrity analysis module 613 is used for intelligent cutting and element expansion based on the parameters of the target display device.
[0064] Please refer to the following: Figure 6A As shown, the layout generation strategy 608 may further include a basic strategy 608a and a decision layer 608b, wherein the basic strategy 608a can be combined with... Figure 6B The layout generation strategy model 610 corresponds to the decision layer 608b, which can be related to... Figure 6B Corresponding to the multimodal semantic analysis system 612 in the layout generation strategy 608, a graph attention network can be used as the decision layer for the layout, which can optimize the layout generation and make the generated layout close to the semantic features of the original visual content.
[0065] This embodiment obtains a semantically aware reinforced layout generator (SARLG) by using reinforcement learning and semantically guided iterative layout generation strategies. Through a cyclical "create-evaluate-optimize" process, the layout generation strategy is self-reinforcing, which improves the ability of the layout generation strategy, increases the efficiency of target layout generation, and ensures that the generated target candidate layouts meet display requirements while maintaining the highest fidelity to the semantic intent and visual narrative of the original visual content.
[0066] This embodiment improves the consistency between the adjusted layout and the intended meaning of the original visual content by comprehensively analyzing the overall semantic context information and hierarchical importance of visual objects during the layout adjustment process.
[0067] Figure 7 The flowcharts of other embodiments of the cross-screen visual content generation method of this disclosure are shown. For example... Figure 7 As shown, the cross-screen visual content generation method may include the following steps: Step 701: Based on semantic analysis of the original visual content, determine the original layout of visual objects and the weight value of each visual object in the original visual content.
[0068] In this embodiment, step 701 and as shown in the figure Figure 2 The steps shown are the same as in step 201; please refer to the original text for the same parts. Figure 2 The corresponding parts of the illustrated implementation are not described in detail here.
[0069] Step 702: Based on the weighted fusion of weight values of multiple dimensions, determine the comprehensive weight value of each visual object, and sort the visual objects in descending order according to the comprehensive weight value.
[0070] In this embodiment, each visual object can have weight values in multiple dimensions. These dimensions may include: positional relationship dimension, semantic relationship dimension, historical data dimension, and distance from the screen edge dimension. The embodiments of this disclosure do not limit this. For example, the weight values of multiple dimensions of a visual object can be determined using the multi-layer attention mechanism of GAT. The weight value S can be determined by calculating relative position, semantic similarity, and historical data. obj S scene S interaction and D screen .
[0071] In this embodiment, the execution entity (e.g., Figure 1 Server 105 can determine the comprehensive weight value of each visual object by weighted fusion of weight values from multiple dimensions, and sort the visual objects in descending order of comprehensive weight value. Specifically, after obtaining the weight values of each visual object across multiple dimensions, the server can label the visual objects with corresponding weight values, then calculate the comprehensive weight value of each visual object through weighted fusion based on the labeled weight values, and finally sort the visual objects in descending order based on the comprehensive weight value.
[0072] In an optional example, such as Figure 8AAs shown, the original visual content 801 includes four visual objects: the mobile phone trademark 802, the mobile phone selling points 803, the front of the mobile phone 804, and the back of the mobile phone 805. These visual objects can be weighted and sorted according to their assigned weight values 806. Specifically, the weight values for the four dimensions of the mobile phone trademark 802 are 0.523, 0.635, 0.702, and 0.388; the weight values for the four dimensions of the mobile phone selling points 803 are 0.455, 0.833, 0.653, and 0.454; the weight values for the four dimensions of the front of the mobile phone 804 are 0.363, 0.521, 0.367, and 0.712; and the weight values for the four dimensions of the back of the mobile phone 805 are 0.345, 0.523, 0.205, and 0.691. The weighted sorting yields the comprehensive weight value W for each visual object. Among them, the mobile phone trademark 802 is the first semantic object in the ranking, with a comprehensive weight value of 0.5144; the mobile phone selling point 803 is the second semantic object in the ranking, with a comprehensive weight value of 0.481; the front of the mobile phone 804 is the third semantic object in the ranking, with a comprehensive weight value of 0.4423; and the back of the mobile phone 805 is the fourth semantic object in the ranking, with a comprehensive weight value of 0.4410.
[0073] Step 703: Based on the parameters of the target display device and the original layout, calculate the loss rate of each visual object displayed on the target display device using an edge collision algorithm.
[0074] In this embodiment, the aforementioned execution entity can calculate the loss rate of each visual object displayed on the target display device using an edge collision algorithm, based on the parameters of the target display device and the original layout. Specifically, when determining the loss rate of each visual object using the edge collision algorithm, the original layout can be directly projected onto the target display device according to its parameters, and the overall scaling ratio can be adjusted to maximize display. The loss rate of each visual object is then calculated using the edge collision algorithm. The parameters of the target display device may include display size, resolution, and irregular screen space matrix, etc.
[0075] In an optional example, such as Figure 8B As shown, based on the screen parameters 808 of the target display device, the overall scaling ratio of the original layout 807 of the original visual content 801 is adjusted to maximize the overall display of the original layout 807 on the target display device. The loss rate 810 of each visual object in the original layout 807 is calculated by the edge collision algorithm 809. It can be found that the loss rate of the mobile phone trademark 802 is 1, the loss rate of the mobile phone selling point 803 is 0.5, the loss rate of the front of the mobile phone 804 is 0.6, and the loss rate of the back of the mobile phone 805 is 0.6.
[0076] Step 704: In response to the fact that the loss rate of a preset number of visual objects sorted in descending order is greater than a preset loss threshold, the original layout is adjusted based on the parameters of the target display device and the comprehensive weight value to determine a target layout that is consistent with the overall semantic features of the original visual content.
[0077] In this embodiment, the execution entity can adjust the original layout based on the parameters of the target display device and the comprehensive weight value in response to a preset number of visual objects in descending order having a loss rate exceeding a preset loss threshold, thereby determining a target layout consistent with the overall semantic features of the original visual content. The preset number of visual objects in the order of priority and the preset loss threshold can be set according to specific application scenarios, and the embodiments disclosed herein do not limit this.
[0078] In an optional example, such as Figure 8B As shown, after obtaining the loss rate 810 for each visual object, it is possible to further determine, based on the comprehensive weight value of each visual object, whether the loss rate of a preset number of high-weight visual objects ranked first in descending order is greater than a preset loss threshold 811. For example, in Figure 8B In this context, all four visual objects are high-weighted visual objects whose loss rate needs to be determined. When the loss rate of any of these high-weighted visual objects exceeds a preset loss threshold, the original layout adjustment 812 is triggered. If the loss rate of no of these high-weighted visual objects exceeds the preset loss threshold, the original layout adjustment is not triggered, and the original visual content can be directly displayed on the target display device.
[0079] Step 705: Based on the target layout, generate new visual content that includes visual objects from the original visual content and is displayed on the target display device.
[0080] In this embodiment, step 705 and as shown Figure 2 The steps shown in step 203 are the same; please refer to the original text for the identical parts. Figure 2 The corresponding parts of the illustrated implementation are not described in detail here.
[0081] In this embodiment, by analyzing the weight values of visual objects and triggering adjustments to the original layout, the display of core objects in the original visual content on the target display device can be guaranteed, ensuring the transmission of the core intent of the original visual content and effectively avoiding the loss of key information or deviation from the intended expression.
[0082] In some optional embodiments of this disclosure, the cross-screen visual content generation method may further include: determining the mask of the occluded area and / or filled area of the new visual content that needs to be repaired based on object semantic recognition of the new visual content; and repairing the occluded area and / or filled area of the new visual content based on the original visual content and the mask of the occluded area and / or filled area.
[0083] In this embodiment, since the original layout of the visual object has been adjusted, the adjusted layout may result in situations where the visual object is occluded and / or the canvas is incomplete due to the expansion of the canvas. Therefore, after generating new visual content, the executing entity (e.g., Figure 1 Server 105 needs to detect the new visual content to determine whether there are any visual objects being occluded and / or the canvas being incomplete due to the expansion of the canvas. If there are any visual objects being occluded and / or the canvas being incomplete due to the expansion of the canvas, the executing entity needs to repair the new visual content.
[0084] In an optional example, such as Figure 9 As shown, when the executing entity determines that new visual content needs to be repaired, it can perform element-level recognition on the new visual content 901 to obtain a mask 902 marking the occluded and filled areas that need to be repaired. The mask 902 includes operable white areas and inoperable black areas; the white areas are the occluded and filled areas that need repair. The executing entity can use the Large Mask (LaMa) algorithm 903 to repair the occluded and filled areas in the new visual content 901 based on the original visual content 904 and the mask 902, ensuring that the occluded and filled areas are displayed as required in the new visual content. Repairing the occluded areas restores the occluded content, while repairing the filled areas fills the incomplete areas of the original visual content 905 caused by canvas expansion (expanded areas), making the canvas of the new visual content 906 complete.
[0085] In some optional embodiments of this disclosure, the cross-screen visual content generation method may further include: extracting content features from the new visual content and extracting style features from the original visual content; constructing transfer features by calculating an attention weight matrix based on the content features and style features; and generating new visual content with a style consistent with the original visual content based on the transfer features.
[0086] In this embodiment, to ensure that the style of the new visual content is consistent with that of the original visual content, the executing entity (e.g., Figure 1 The server (105) can extract the style features of the original visual content and transfer the style features of the original visual content to the new visual content.
[0087] In an optional example, the agent can inject the style features of the original visual content into the new visual content using the Adaptive Attention Normalization (AdaAttN) style transfer attention mechanism. Specifically, the agent can extract the style features F from the original visual content.s Content characteristics of new visual content F c Through attention weight matrix The calculation constructs the transfer feature F cs According to the migration feature F cs Generate new visual content that maintains the same style as the original visual content.
[0088] The cross-screen visual content generation method provided in this disclosure will be described below with reference to specific application scenarios.
[0089] Figure 10 This paper illustrates a first application scenario of the cross-screen visual content generation method of this disclosure, which implements the adjustment of the composition of image content. Specifically, this application scenario uses the cross-screen visual content generation method provided by this disclosure to rearrange the elements (visual objects) in the image (original visual content) according to the original display style and theme after understanding the image intent, for images with significantly different display ratios, so as to facilitate display on the device. Figure 10 As shown, semantic analysis is first performed on the original image 1001 to obtain the original layout 1002 of the original image 1001 and the weight values of each component element in the original image 1001. Then, the original layout 1002 is adjusted according to the parameters of the target display device and the weight values of each component element to obtain a target layout 1003 that is consistent with the overall semantic features of the original image 1001. Finally, a new image 1004 containing each component element of the original image 1001 is generated for display on the target display device based on the target layout 1003. The semantic analysis may include image recognition, element extraction, and intent recognition, where the extracted element is 1005 and the result of intent recognition is 1006.
[0090] Figure 11 This paper illustrates a second application scenario using the cross-screen visual content generation method of this disclosure. This application scenario implements adjustments to image display based on irregularly shaped screens. Specifically, for irregularly shaped screens, such as punch-hole screens, notch screens, and waterdrop screens on mobile phones, the non-displayable punch-hole area may miss some important information when displaying images or videos, resulting in a poor user experience. Therefore, it is necessary to determine the information in this area during display. To better showcase the main theme, the cross-screen visual content generation method provided by this disclosure can be used to appropriately adjust the position of elements (visual objects) in this area (punch-hole area). Specific implementation steps can be found in the descriptions of the preceding implementation examples. Figure 11 As shown, the original image is 1101, and the new image after layout adjustment using the cross-screen visual content generation method of this disclosure is 1102.
[0091] Figure 12This paper illustrates a third application scenario using the cross-screen visual content generation method of this disclosure. This application scenario enables adaptive adjustment of variable screens (foldable / extendable screens). Specifically, for variable-size screens, fixed-size images or videos cannot be displayed well across varying screen sizes. Using the cross-screen visual content generation method provided by this disclosure, images can be dynamically adjusted to adapt to the information display needs of different modes, ensuring a good display effect. Specific implementation steps can be found in the descriptions of the preceding implementation examples. Figure 12 As shown, when the original image is 1201, the new image after layout adjustment using the cross-screen visual content generation method of this disclosure is 1202; when the original image is 1202, the new image after layout adjustment using the cross-screen visual content generation method of this disclosure is 1201.
[0092] Figure 13 A fourth application scenario using the cross-screen visual content generation method of this disclosure is illustrated. This application scenario implements adaptive image cropping. Specifically, for images with different display ratios, after understanding the image intent, the image's fill effect is insufficient to highlight the image's main purpose. Using the cross-screen visual content generation method provided by this disclosure, irrelevant content can be cropped out for display on the device. Specific implementation steps can be found in the descriptions of the preceding implementation examples. Figure 13 As shown, the original image is 1301, 1302 shows the area cropped from the original image, and the new image generated after layout adjustment using the cross-screen visual content generation method disclosed herein is 1303.
[0093] The cross-screen visual content generation method of this disclosure can be applied to an integrated content management platform ecosystem. In this ecosystem, the platform manages thousands of different display configurations and parameter specifications, including large outdoor displays, modular video walls with diverse multi-screen configurations, custom-shaped displays, and digital signage systems. By dynamically adjusting the visual content according to the parameters of the target display screen using the cross-screen visual content generation method of this disclosure, significant commercial value and operational efficiency can be achieved.
[0094] The cross-screen visual content generation method provided by the embodiments of this disclosure simplifies content deployment across heterogeneous display environments, enabling users to seamlessly publish content without considering technical issues related to display parameter compatibility. A simplified one-click publishing workflow and automatic optimization of display modes based on specific screen parameters significantly improve user experience and operational efficiency.
[0095] This technology can maintain absolute visual integrity throughout the entire adaptation process. Figure 1Consistency is a key requirement for business communications and advertising applications, where message integrity and brand consistency are crucial for campaign effectiveness and audience engagement.
[0096] Existing integrated content management platforms aim to provide unified services for thousands of heterogeneous display devices with varying specifications and forms. Applying the methods provided in this disclosure to an integrated content management platform will enable powerful content publishing capabilities across the entire device ecosystem. By implementing a unified "edit once, deploy in batches" function, significant strategic business value will be created for enterprises, streamlining the content creation process and enabling simultaneous deployment across multiple devices, thereby greatly improving operational efficiency and shortening the time to market for digital signage marketing campaigns.
[0097] The method disclosed herein provides a way to generate multi-format visual content with different aspect ratios by comprehensively analyzing visual objects, style attributes, and thematic components in the original visual content, using advanced generative AI while strictly adhering to the original design aesthetics. This intelligent approach significantly optimizes the workflow efficiency of creative professionals and content production teams.
[0098] The method provided in the embodiments of this disclosure has a proprietary object extraction and reconstruction process that ensures all generated visual objects are completely derived from the original visual objects, maintains complete visual consistency, and eliminates the differences between the original visual objects and the output visual objects.
[0099] The methods provided in this disclosure are based on an advanced image generation framework, which extends the system's capabilities to generate visual content in various formats and specifications, enabling seamless cross-platform deployment. Intelligent scaling algorithms dynamically adjust content presentation for variable display environments, ensuring optimal visual fidelity across different screen configurations and device specifications.
[0100] Further reference Figure 14 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of a cross-screen visual content generation device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0101] like Figure 14 As shown, the cross-screen visual content generation device in this embodiment may include: a semantic analysis module 1401, a layout adjustment module 1402, and a content generation module 1403.
[0102] The semantic analysis module 1401 can be configured to perform semantic analysis on the original visual content to determine the original layout of visual objects and the weight values of each visual object in the original visual content. The layout adjustment module 1402 can be configured to adjust the original layout based on the parameters and weight values of the target display device to determine a target layout consistent with the overall semantic features of the original visual content. The content generation module 1403 can be configured to generate new visual content containing the visual objects from the original visual content for display on the target display device, based on the target layout.
[0103] In some alternative implementations of this disclosure, the semantic analysis module 1401 may include: The instance segmentation unit is configured to perform instance segmentation on the original visual content and determine the category, bounding box, and mask of the visual objects in the original visual content. The scene semantic analysis unit is configured to perform scene semantic analysis on the original visual content based on the category, bounding box and mask of the visual object, and determine the scene semantic features. The layout and weight determination unit is configured to determine the original layout of the visual objects and the weight values of each visual object based on the category, bounding box, and mask of the visual objects and the semantic features of the scene.
[0104] In some optional implementations of this disclosure, the scene semantic analysis unit may include: The scene label generation subunit is configured to extract image features and text features from the category, bounding box and mask of the visual object, perform feature alignment, calculate the similarity after feature alignment, and determine the scene label of the original visual content. The scene semantic reasoning subunit is configured to perform scene semantic reasoning on the scene tags based on the knowledge graph to determine the scene semantic features of the original visual content.
[0105] In some alternative implementations of this disclosure, the layout and weight determination unit may include: The layout determination subunit is configured to determine the original layout of the visual object based on the category, bounding box, and mask of the visual object; The weight determination subunit is configured to construct a visual object relationship graph based on the category, bounding box, and mask of the visual object and the semantic features of the scene, and to determine the weight value of each visual object based on the visual object relationship graph.
[0106] In some alternative implementations of this disclosure, the layout adjustment module 1402 may include: The layout adjustment unit is configured to adjust the original layout according to the parameters of the target display device and the weight value through a layout generation strategy, and generate a candidate layout group containing multiple candidate layouts. The layout filtering unit is configured to perform semantic analysis on the candidate layouts and compare them with the overall semantic features of the original visual content to determine a plurality of target candidate layouts whose differences from the overall semantic features are less than a preset difference threshold, and to determine the target layout from the plurality of target candidate layouts.
[0107] In some alternative implementations of this disclosure, the layout adjustment module 1402 may further include: The strategy update unit is configured to update the layout generation strategy based on the target candidate layout; The layout adjustment unit is further configured to adjust the original layout according to the parameters of the target display device and the weight value through the updated layout generation strategy, and generate a candidate layout group containing multiple candidate layouts.
[0108] In some optional implementations of this disclosure, the semantic analysis module 1401 is further configured to perform semantic analysis on the original visual content through a multimodal semantic analysis system to determine the original layout of visual objects in the original visual content and the weight values of each visual object. The layout filtering unit is further configured to input the candidate layout group into the multimodal semantic analysis system, perform semantic analysis on the candidate layouts and compare them with the overall semantic features of the original visual content, and determine a number of target candidate layouts whose differences from the overall semantic features are less than a preset difference threshold.
[0109] In some alternative implementations of this disclosure, each visual object has weight values in multiple dimensions; The cross-screen visual content generation device also includes: The weighted sorting module is configured to determine the comprehensive weight value of each visual object based on the weight values of the multiple dimensions, and sort the visual objects in descending order according to the comprehensive weight value. The loss calculation module is configured to calculate the loss rate of each visual object displayed on the target display device based on the parameters of the target display device and the original layout, using an edge collision algorithm. The layout adjustment module 1402 is further configured to adjust the original layout based on the parameters of the target display device and the comprehensive weight value in response to the loss rate of the preset number of visual objects in descending order being greater than a preset loss threshold, and to determine a target layout that is consistent with the overall semantic features of the original visual content.
[0110] In some alternative implementations of this disclosure, the cross-screen visual content generation apparatus further includes: The occlusion / expansion repair module is configured to determine the mask of the occluded area and / or filled area of the new visual content that needs to be repaired based on object semantic recognition of the new visual content; and repair the occluded area and / or filled area of the new visual content based on the original visual content and the mask of the occluded area and / or filled area.
[0111] In some alternative implementations of this disclosure, the cross-screen visual content generation apparatus further includes: The style injection module is configured to extract content features from the new visual content and style features from the original visual content; based on the content features and the style features, a transfer feature is constructed by calculating an attention weight matrix; based on the transfer feature, new visual content with the same style as the original visual content is generated.
[0112] According to embodiments of this disclosure, this disclosure also provides an electronic device, a computer-readable storage medium, and a computer program product.
[0113] Figure 15 This is a block diagram of an electronic device suitable for implementing embodiments of the present disclosure. For example... Figure 15 As shown, the electronic device includes one or more processors 1501, a memory 1502, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 15 Take a processor 1501 as an example.
[0114] The memory 1502 is the non-transitory computer-readable storage medium provided in this disclosure. The memory stores instructions executable by at least one processor to cause the at least one processor to perform the cross-screen visual content generation method provided in this disclosure. The non-transitory computer-readable storage medium of this disclosure stores computer instructions for causing a computer to perform the cross-screen visual content generation method provided in this disclosure.
[0115] Memory 1502, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the cross-screen visual content generation method in this disclosure embodiment (e.g., appendix). Figure 14 The semantic analysis module 1401, layout adjustment module 1402, and content generation module 1403 are shown. The processor 1501 executes various server functions and data processing by running non-transient software programs, instructions, and modules stored in the memory 1502, thereby realizing the cross-screen visual content generation method in the above method embodiments.
[0116] Memory 1502 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the use of the electronic device during historical video playback. Furthermore, memory 1502 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 1502 may optionally include memory remotely located relative to processor 1501, and these remote memories may be connected via a network to the electronic device performing the cross-screen visual content generation method. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0117] The electronic device for the cross-screen visual content generation method may further include an input device 1503 and an output device 1504. The processor 1501, memory 1502, input device 1503, and output device 1504 can be connected via a bus or other means. Figure 15 Taking the example of a connection between China and Israel via a bus.
[0118] Input device 1503 can receive input digital or character information, as well as key signal input related to user settings and function control of electronic devices playing video streams, such as touch screens, keypads, mice, trackpads, touchpads, joysticks, one or more mouse buttons, trackballs, joysticks, etc. Output device 1504 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The display device may include, but is not limited to, liquid crystal displays (LCDs), light-emitting diode (LED) displays, and plasma displays. In some embodiments, the display device may be a touch screen.
[0119] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, application-specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0120] These computational programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0121] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0122] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0123] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0124] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0125] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A method for generating cross-screen visual content, comprising: Based on semantic analysis of the original visual content, the original layout of visual objects and the weight value of each visual object in the original visual content are determined. The original layout is adjusted based on the parameters of the target display device and the weight values to determine a target layout that is consistent with the overall semantic features of the original visual content. Based on the target layout, new visual content containing visual objects from the original visual content is generated and displayed on the target display device.
2. The method according to claim 1, wherein, The step of determining the original layout of visual objects and the weight value of each visual object in the original visual content based on semantic analysis includes: The original visual content is segmented into instances to determine the category, bounding box, and mask of the visual objects in the original visual content. Based on the category, bounding box, and mask of the visual object, scene semantic analysis is performed on the original visual content to determine scene semantic features; Based on the category, bounding box, and mask of the visual object and the semantic features of the scene, the original layout of the visual object and the weight value of each visual object are determined.
3. The method according to claim 2, wherein, The step of performing scene semantic analysis on the original visual content based on the category, bounding box, and mask of the visual object to determine scene semantic features includes: For the category, bounding box, and mask of the visual object, extract image features and text features, perform feature alignment, calculate the similarity after feature alignment, and determine the scene label of the original visual content; Based on the knowledge graph, scene semantic reasoning is performed on the scene tags to determine the scene semantic features of the original visual content.
4. The method according to claim 2, wherein, The process of determining the original layout of the visual objects and the weight values of each visual object based on the category, bounding box, and mask of the visual objects and the semantic features of the scene includes: The original layout of the visual object is determined based on its category, bounding box, and mask. Based on the category, bounding box, and mask of the visual object and the semantic features of the scene, a visual object relationship graph is constructed, and the weight value of each visual object is determined based on the visual object relationship graph.
5. The method according to claim 1, wherein, The step of adjusting the original layout based on the parameters of the target display device and the weight values to determine a target layout consistent with the overall semantic features of the original visual content includes: The original layout is adjusted according to the parameters of the target display device and the weight value through the layout generation strategy to generate a candidate layout group containing multiple candidate layouts; Based on semantic analysis of the candidate layouts and comparison with the overall semantic features of the original visual content, a number of target candidate layouts with differences from the overall semantic features less than a preset difference threshold are determined, and the target layout is determined from the number of target candidate layouts.
6. The method according to claim 5, wherein, The step of adjusting the original layout based on the parameters of the target display device and the weight values to determine a target layout consistent with the overall semantic features of the original visual content further includes: Based on the target candidate layout, the layout generation strategy is updated; By using the updated layout generation strategy, the original layout is adjusted according to the parameters of the target display device and the weight value to generate a candidate layout group containing multiple candidate layouts.
7. The method according to claim 5, wherein, The step of determining the original layout of visual objects and the weight value of each visual object in the original visual content based on semantic analysis includes: The original visual content is semantically analyzed by a multimodal semantic analysis system to determine the original layout of visual objects and the weight value of each visual object in the original visual content. The step of determining multiple target candidate layouts whose differences from the overall semantic features of the original visual content are less than a preset difference threshold based on semantic analysis of the candidate layouts includes: The candidate layout group is input into the multimodal semantic analysis system, the candidate layout is semantically analyzed and compared with the overall semantic features of the original visual content, and multiple target candidate layouts whose differences from the overall semantic features are less than a preset difference threshold are determined.
8. The method according to any one of claims 1-7, wherein, Each visual object has weight values in multiple dimensions; The method further includes: Based on the weighted fusion of the weight values of the multiple dimensions, a comprehensive weight value for each visual object is determined, and the visual objects are sorted in descending order according to the comprehensive weight value. Based on the parameters of the target display device and the original layout, the loss rate of each visual object displayed on the target display device is calculated using an edge collision algorithm; The step of adjusting the original layout based on the parameters of the target display device and the weight values to determine a target layout consistent with the overall semantic features of the original visual content includes: In response to the fact that the loss rate of the preset number of visual objects in descending order is greater than a preset loss threshold, the original layout is adjusted based on the parameters of the target display device and the comprehensive weight value to determine a target layout that is consistent with the overall semantic features of the original visual content.
9. The method according to any one of claims 1-7, further comprising: Based on object semantic recognition of the new visual content, the mask of the occluded area and / or filled area that needs to be repaired in the new visual content is determined. Based on the original visual content and the mask of the occluded and / or filled areas, the occluded and / or filled areas of the new visual content are repaired.
10. The method according to any one of claims 1-7, further comprising: Extract content features from the new visual content and style features from the original visual content; Based on the content features and style features, transfer features are constructed by calculating the attention weight matrix; Based on the migration features, new visual content with the same style as the original visual content is generated.
11. A cross-screen visual content generation device, comprising: The semantic analysis module is configured to determine the original layout of visual objects and the weight value of each visual object in the original visual content based on semantic analysis of the original visual content. The layout adjustment module is configured to adjust the original layout based on the parameters of the target display device and the weight value, and determine a target layout that is consistent with the overall semantic features of the original visual content. The content generation module is configured to generate new visual content, which includes visual objects from the original visual content, for display on the target display device, based on the target layout.
12. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores information that can be executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-10.
13. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-10.
14. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1-10.