Data processing method and device, equipment and storage medium
By performing semantic segmentation on images and videos processed by smart glasses or the cloud, creative output information is generated, solving the problem of insufficient interactivity in existing smart glasses technology and improving the user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUIZHOU TCL CLOUD INTERNET CORP TECH CO LTD
- Filing Date
- 2026-01-04
- Publication Date
- 2026-05-15
AI Technical Summary
Existing smart glasses technology mainly focuses on first-person perspective shooting and information overlay, lacking creative output and resulting in a lack of interactivity.
By performing semantic segmentation on images and videos to be processed, semantic segmentation entity and element labels are generated. Based on these entities and labels, creative output information, such as story text, comics or videos, is generated and processed and projected using smart glasses or the cloud.
It enables creative output of real-world scenarios, enhances user experience and interactivity, and provides creative interactive scenarios.
Smart Images

Figure CN122053932A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer vision, specifically to a data processing method, apparatus, device, and storage medium. Background Technology
[0002] Current smart glasses technology primarily focuses on first-person perspective shooting, providing information overlay (such as navigation data visualization) and artificial intelligence (AI) visual assistance, mainly for "information transmission" or "recording," such as generating navigation images, generating translated text based on voice input, or generating text records based on voice input. However, this approach fails to provide creative output, resulting in a lack of interactivity in smart glasses. Summary of the Invention
[0003] This application provides a data processing method, apparatus, computer device, and computer-readable storage medium for converting image information into creative output information in real time, thereby providing creative interactive scenarios and enhancing user experience.
[0004] The technical solution adopted by this invention to solve the problem is as follows: In a first aspect, this application provides a data processing method, comprising: acquiring an image to be processed and / or a video to be processed; performing semantic segmentation on the image to be processed and / or the video to be processed to obtain semantic segmentation entities and element tags corresponding to the semantic segmentation entities, wherein the element tags are used to characterize the narrative elements corresponding to each semantic segmentation entity; and generating playback information based on the semantic segmentation entities and the element tags corresponding to the semantic segmentation entities. In some implementations of this application, generating playback information based on the semantic segmentation entity and the corresponding element tag includes: A real-time narrative analysis framework is generated based on the semantic segmentation entity, the element tags corresponding to the semantic segmentation entity, and the preset narrative template. This real-time narrative analysis framework is used to characterize the narrative logic of the information to be played. The target processing model is invoked to process the real-time narrative analysis framework to obtain the information to be played.
[0005] In some implementations of this application, the target processing model is invoked to process the real-time narrative analysis framework to obtain the information to be played, including: The first processing model is invoked to generate text for the real-time narrative analysis framework to obtain the story text; And / or, The second processing model is invoked to generate images from the real-time narrative analysis framework to obtain comics; And / or, The third processing model is invoked to generate video from the real-time narrative analysis framework to obtain the video. The story text, the comic, and the video are the information to be played, and the first processing model, the second processing model, and the third processing model are included in the target processing model.
[0006] In some embodiments of this application, semantic segmentation is performed on the image and / or video to be processed to obtain semantic segmentation entities and corresponding element labels, including: The fourth processing model is invoked to perform semantic segmentation on the image and / or video to be processed, so as to obtain the semantic segmentation entity and the corresponding feature label assigned to the semantic entity. The fourth processing model is a semantic segmentation model trained in combination with the feature label.
[0007] In some embodiments of this application, obtaining the image and / or video to be processed includes: The image and / or video to be processed are acquired using the camera of a wearable device, including but not limited to smart glasses and smart helmets.
[0008] In some embodiments of this application, after generating playback information based on the semantic segmentation entity and the corresponding element tags, the method further includes: When the information to be played is story text, multi-character voice synthesis is performed on the story script to obtain a voice script, and then the voice script is played. When the information to be played is a comic, the comic is overlaid and projected using virtual reality. When the information to be played is a video, a virtual reality overlay projection is applied to the video.
[0009] In some embodiments of this application, after generating playback information based on the semantic segmentation entity and the corresponding element tags, the method further includes: The playback operation of the information to be played is performed based on gestures and / or eye-tracking focus recognition in order to obtain feedback information; The factual narrative analysis framework was adjusted based on this feedback.
[0010] Secondly, this application provides a data processing apparatus, including: an acquisition module for acquiring images and / or videos to be processed; The processing module is used to perform semantic segmentation on the image and / or video to be processed, so as to obtain semantic segmentation entities and the element labels corresponding to the semantic segmentation entities. The element labels are used to characterize the narrative elements corresponding to each semantic segmentation entity. The generation module is used to generate playback information based on the semantic segmentation entity and the corresponding element label.
[0011] Thirdly, this application also provides a computer device, which includes: One or more processors; Memory; and One or more applications, wherein the applications are stored in memory and configured to be executed by a processor to implement the data processing method of any of the first aspects.
[0012] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps of the data processing method of any one of the first aspects.
[0013] The beneficial effects of this invention are as follows: semantic segmentation processing is performed on the acquired images and / or videos to be processed to obtain semantic segmentation entities and corresponding element labels; finally, creative output information is generated based on the semantic segmentation entities and corresponding element labels, thereby realizing creative output of real-world scenes, providing creative interactive scenarios, and increasing user experience. Attached Figure Description
[0014] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments recorded in the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0015] Figure 1 This is a schematic diagram of the system architecture provided in an embodiment of the present invention; Figure 2 This is a schematic flowchart of an embodiment of the data processing method provided by the present invention; Figure 3 This is a flowchart illustrating a data processing method for smart glasses provided in an embodiment of the present invention; Figure 4 This is a schematic diagram of a specific embodiment of the data processing device provided in this invention. Figure 5 This is a schematic diagram of an embodiment of the computer device provided in this invention. Detailed Implementation
[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.
[0017] In the description of this application, the terms "first," "second," "third," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first," "second," "third," etc., may explicitly or implicitly include one or more features.
[0018] In this application, the term "exemplary" is used to mean "used as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use this application. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that this application can be made without using these specific details. In other instances, well-known structures and processes are not described in detail to avoid obscuring the description of this application with unnecessary detail. Therefore, this application is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0019] It should be noted that since the method in this application embodiment is executed in a computer device, the processing objects of each computer device exist in the form of data or information, such as time, which is essentially time information. It is understood that if size, quantity, position, etc. are mentioned in subsequent embodiments, they are all corresponding data that exist so that the computer device can process them. Specific details will not be elaborated here.
[0020] Current smart glasses technology mainly focuses on first-person perspective shooting, providing information overlay (such as navigation data visualization) and AI visual assistance, primarily for "information transmission" or "recording," such as generating navigation images, generating translated text based on voice input, or generating text records based on voice input. However, this cannot provide creative output, resulting in a lack of interactivity in smart glasses.
[0021] To address this technical problem, this application provides the following technical solution: acquiring an image and / or video to be processed; performing semantic segmentation on the image and / or video to obtain semantic segmentation entities and corresponding element tags, whereby the element tags characterize the narrative elements corresponding to each semantic segmentation entity; and generating playback information based on the semantic segmentation entities and their corresponding element tags. This approach involves semantic segmentation of the acquired image and / or video to obtain semantic segmentation entities and their corresponding element tags; and finally, generating creative output information based on these semantic segmentation entities and their corresponding element tags. This achieves creative output of real-world scenarios, thereby providing creative interactive scenarios and enhancing the user experience.
[0022] For ease of understanding, the relevant technologies involved in this application are described below.
[0023] CLOAST is a real-time narrative analysis framework consisting of six core elements: Character, Location, Object, Situation, Act, and Theme. It is used to transform visual input into structured story data.
[0024] Augmented Reality (AR): A technology that overlays virtual information (images, text, etc.) onto the real world.
[0025] This application provides a data processing method, apparatus, device, and storage medium for converting image scenes into creative output information in real time, thereby providing creative interactive scenarios and enhancing user experience. The electronic device provided in this application can be implemented as various types of user terminals or as a server.
[0026] Electronic devices, by running the data processing method provided in the embodiments of this application, can transform image scenes into creative output information in real time, thereby providing creative interactive scenarios and enhancing the user experience.
[0027] The above methods can be applied to many content creation fields, such as converting images captured in real time by smart glasses into new types of playable information; or converting videos shot by mobile phones into new types of playable information.
[0028] In one exemplary solution, this data processing method can be applied to content interaction in smart glasses. For instance, in a smart glasses usage scenario, users can generate story text, comics, or videos with plots from captured images in real time. The specific process can be as follows: When wearing smart glasses, the user captures the current real-world scene to obtain a video or image. After activating the content creation function, the smart glasses perform semantic segmentation processing on the video or image to obtain semantic segmented entities and their corresponding element tags. Finally, based on the semantic segmented entities and their corresponding element tags, information to be played is generated. It should be understood that, when applied to smart glasses scenarios, due to the limited processing power of smart glasses, the processing of the video or image can also be applied to the cloud.
[0029] It should be understood that the above is only an example of a data processing application scenario. There are many other possible application scenarios, which are not limited here.
[0030] The data processing method provided in this application embodiment is applied to, for example, Figure 1 The system architecture diagram shown is for your reference. Figure 1To support a data processing method, the terminal device 100 connects to the server 300 via a network 200, and the server 300 connects to the database 400. The network 200 can be a wide area network (WAN), a local area network (LAN), or a combination of both. The client for implementing the content operation plan is deployed on the terminal device 100, or it can run on the terminal device 100 as a standalone application. The specific form of the client is not limited here. The server 300 involved in this application can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms. The terminal device 100 can be a device that includes both receiving and transmitting hardware, i.e., a device with receiving and transmitting hardware capable of performing bidirectional communication on a bidirectional communication link. Such a device can include cellular or other communication devices with single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays. The specific terminal device 100 can be a desktop terminal or a mobile terminal. Specifically, it can also be one of the following: Augmented Reality (AR) devices, mobile phones, tablets, laptops, in-vehicle devices, wearable devices, smart TVs, smart home appliances, aircraft, or intelligent voice interaction devices. The solution provided in this application can be completed by the terminal device 100 and the server 300 working together. The database 400, in short, can be viewed as an electronic filing cabinet—a place to store electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, shared by multiple users, with minimal redundancy, and independent of applications. A Database Management System (DBMS) is a computer software system designed for managing databases, generally possessing basic functions such as storage, retrieval, security, and backup. Database management systems (DBMS) can be categorized based on the database model they support, such as relational or Extensible Markup Language (XML); or based on the type of computer they support, such as server clusters or mobile devices; or based on the query language used, such as Structured Query Language (SQL) or XQuery; or based on performance priorities, such as maximum scale or highest operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, by simultaneously supporting multiple query languages.In this application, the database 400 can be used to store data and corresponding processing models, such as images to be processed, videos to be processed, and information to be played.
[0031] Those skilled in the art will understand that Figure 1 The system architecture diagram shown is one possible system architecture for this application and does not constitute a limitation on the system architecture of this application. Other system architectures may include more advanced architectures. Figure 1 The number of more or fewer terminal devices or servers shown, for example Figure 1 Only one server is shown in the diagram. It is understood that the system architecture may also include one or more other terminal devices or servers, which are not limited here.
[0032] It should be noted that, Figure 1 The system architecture shown is an example. The servers and scenarios described in the embodiments of this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of servers and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0033] like Figure 2 The diagram shown is a flowchart of an embodiment of the data processing method in this application. The following description, using a server as the execution entity, details the data processing method, which may include the following steps 201-203: 201. Obtain the images and / or videos to be processed.
[0034] In this embodiment, the image and / or video to be processed can be real-time captured data or historical captured data; the specific type is not limited here. That is, the server can receive the image and / or video to be processed sent by the real-time capturing device, or it can directly receive the image and / or video to be processed downloaded and uploaded by the user from a third-party platform.
[0035] The real-time shooting device can be one of the following: a user's mobile phone, tablet, smart glasses, smart helmet, in-vehicle equipment, wearable device, smart TV, smart home appliance, aircraft, or smart voice interaction device.
[0036] Optionally, the image and / or video to be processed may also carry a timestamp, so that the timeline of the generated playback information can be determined based on the timestamp during subsequent processing.
[0037] 202. Perform semantic segmentation on the image and / or video to be processed to obtain semantic segmentation entities and corresponding element labels. The element labels are used to characterize the narrative elements corresponding to each semantic segmentation entity.
[0038] After receiving the image and / or video to be processed, the server can call the fourth processing model to perform semantic segmentation on the image and / or video to obtain the corresponding semantic segmentation entities and the element labels corresponding to each semantic segmentation entity on the image and / or video.
[0039] Optionally, the element label can be an element in the CLOAST framework, that is, the element label can include Character, Location, Object, Situation, Act, and Theme.
[0040] In this embodiment, to improve the processing efficiency of semantic segmentation, the server can train a semantic segmentation model that meets the requirements of this embodiment. That is, the fourth processing model is obtained by training in conjunction with feature labels.
[0041] In one exemplary scheme, the training process of the semantic segmentation model (i.e., the fourth processing model) can be as follows: First, a labeled dataset (i.e., a training dataset) is created. In this process, existing large-scale semantic segmentation datasets (such as COCO, ADE20K, Cityscapes, etc.) can be used. These datasets already contain pixel-level annotations of objects (people, cars, buildings, trees, etc.) in images. In this embodiment, secondary or extended annotations are also needed for these datasets. In the secondary or extended annotations, feature labels are added. For example, for a "portrait" in an image, after labeling "This is a 'person'," a narrative label is added. For example: "Person -> Character," "Buildings, Streets, Houses - Scene -> Location," "Books, Cars, Phones -> Objects."
[0042] During the annotation process, complex tags can be manually defined. For example, by analyzing image sequences, actions such as walking, running, and talking can be labeled and categorized as action (Act) element tags. By analyzing the overall atmosphere (such as dim lighting or vibrant colors), theme (such as suspense or romance) element tags can be labeled.
[0043] Then, a customized semantic segmentation model is trained. In this process, a lightweight semantic segmentation model (such as MobileNetV3-based DeepLab, BiSeNet, etc.) is trained using the dataset annotated with CLOSAT enhancements mentioned above, so that it can run in real time on smart glasses.
[0044] During training, the training objective is set to enable the model to directly output a "pixel mask with CLOSAT labels" after receiving an image. This means that in the output image, the value of each pixel no longer represents "car" or "person", but directly represents "object" or "character".
[0045] Through the above training process, the fourth processing model can be obtained. Then, in online applications, the image and / or video to be processed are used as the fourth processing model. This model can output a pixel mask of the same size as the original image. In this pixel mask, all pixels belonging to characters can be labeled as 1, all pixels belonging to the scene can be labeled as 2, all pixels belonging to objects can be labeled as 3, and so on. It should be understood that the 1, 2, and 3 mentioned above refer to different labels; this is merely an exemplary scheme, and no specific limitation is made here.
[0046] It should be understood that when the input data is a video to be processed, the video needs to be preprocessed to generate an image sequence.
[0047] Optionally, to facilitate subsequent processing, in this embodiment, after the server obtains the semantic segmentation entity and the corresponding feature label, it can further process the above information to obtain structured data. In an exemplary scheme, the server identifies connected pixel blocks with the same label (e.g., a whole region representing a Character); extracts the information of these pixel blocks, and integrates them into a structured JSON object.
[0048] 203. Generate playback information based on the semantic segmentation entity and the element tags corresponding to the semantic segmentation entity.
[0049] After obtaining the semantic segmentation entity and the corresponding element tag, the server generates a real-time narrative analysis framework based on the semantic segmentation entity, the corresponding element tag, and a preset narrative template. This real-time narrative analysis framework is used to represent the narrative logic of the information to be played. The target processing model is then called to process the real-time narrative analysis framework to obtain the information to be played.
[0050] That is, the server converts the semantic segmentation entity and the corresponding element label into AI parameters, and builds a real-time narrative analysis framework based on the preset narrative template and the AI parameters. Finally, the target processing model is called to generate and process the data in the real-time narrative analysis framework to obtain the information to be played.
[0051] Optionally, the process by which the server converts the semantic segmentation entity and the corresponding feature label into AI parameters may include at least one of the following processes: In one exemplary scheme, the server calculates the context and context tension score based on the semantic segmentation entity and its corresponding feature label. During this process, the parameters of each semantic segmentation entity identified as a Character may include: Character ID: A unique ID used to track the same character (e.g., character 1, character 2) in consecutive frames.
[0052] Bounding Box: A rectangular coordinate system [x, y, width, height] that precisely defines the position and size of the character in the current frame.
[0053] Pose Keypoints: Body joints (head, shoulders, elbows, knees, etc.) extracted using pose estimation algorithms (such as Open Pose). These are crucial for determining the character's orientation.
[0054] Timestamp: The time of the current frame, used to analyze dynamic changes.
[0055] After obtaining the aforementioned data through the semantic segmentation entity and the corresponding feature label, the server will perform a series of calculations to transform the spatial relationship into quantifiable parameters.
[0056] In one possible implementation, the server calculates the proximity of two characters, that is, the narrative parameter of the physical distance between two entities labeled "character". In this case, the server can calculate the Euclidean distance between the center points of the character bounding boxes; then, referring to interpersonal distance theory, it maps this distance to different narrative categories: Intimate distance (i.e., adjacent distance less than 0.5 meters): If the calculated distance falls within this range, the server can assign a high "intimacy score" or "tension score" to the two entities. This could be interpreted as: hugging, whispering, conflict, or confrontation.
[0057] Personal distance (i.e., adjacent distance between 0.5 and 1.2 meters): the normal conversational distance between friends. The server can assign a neutral "social score" to these two entities. This can be interpreted as: dialogue and interaction.
[0058] Social distance (i.e., adjacent distance between 1.2 and 3.6 meters): Formal interaction distance in non-personal matters. The server can assign a high "alienation score" to these two entities. Interpreted as: unfamiliarity, observation, formal meeting.
[0059] Public distance (i.e., adjacent distance greater than 3.6 meters): almost no direct interaction. Interpreted as: irrelevant, part of the environment.
[0060] Furthermore, the server can also calculate the relative orientation of two characters, allowing for the interpretation of their relationship based on this orientation. In real-world scenarios, the narrative meaning of two people standing close together can be drastically different depending on whether they are back-to-back or face-to-face. Therefore, relative orientation can more accurately describe the contextual narrative between two characters. This calculation requires the use of pose keypoint data.
[0061] The interpretation of the relative orientation can include the following: Face to face: Judged by key points on the head. If two characters are close together and facing each other, the "tension score" will rise sharply. This could mean a heated conversation, an argument, or a romantic gaze.
[0062] Back to back: indicates separation, contradiction, and ignoring each other.
[0063] Side by side: indicates walking together, cooperating, or looking at a certain place together.
[0064] One in front and one behind: If combined with posture, one person is facing forward and the other is looking back at the one in front, this may be interpreted as tracking, chasing or protecting.
[0065] Furthermore, the server can also calculate the dynamic change rate based on consecutive frames and interpret the action descriptions and contextual descriptions between two characters based on the dynamic change rate. Specifically, this can include the following: Distance decreasing rapidly: Calculates the rate at which the distance decreases. If this value is negative and has a large absolute value, the system will interpret it as "approaching," "rushing towards," or "chasing."
[0066] Distance increases rapidly: If the value is positive and the absolute value is large, it is interpreted as "leaving", "escaping" or "separating".
[0067] Maintaining a stable distance while moving: If the relative positions of two characters remain unchanged, but their overall position in the frame moves, this is interpreted as "moving together".
[0068] In this process, the final output may include: Context: a text label, such as "standoff," "chase," or "secret conversation." Context tension score: a numerical value from 0 to 100.
[0069] In another exemplary approach, the server predicts the topic based on the semantic segmentation entity and the corresponding element label. This process can integrate all elements for analysis. For example, given a scene of "bar," an item of "glass," and an overall color scheme of "dark tones," the topic can be inferred to be "detective reasoning."
[0070] If the scene is a "park", the characters include two people, and the situation is intimate, then the theme can be inferred to be a "romantic movie".
[0071] The server can refer to the above process to perform parameter transformation on the identified semantic segmentation entities and their element labels, and construct a real-time narrative analysis framework. Then, the server calls the target processing model to generate the information to be played based on this real-time narrative analysis framework; alternatively, the server calls the target processing model to generate the information to be played based on the real-time narrative analysis framework. The specific process can be as follows: The first processing model is invoked to generate text for the real-time narrative analysis framework to obtain the story text; And / or, The second processing model is invoked to generate images from the real-time narrative analysis framework to obtain comics; And / or, The third processing model is invoked to generate video from the real-time narrative analysis framework to obtain the video. The story text, the comic, and the video are the information to be played, and the first processing model, the second processing model, and the third processing model are included in the target processing model.
[0072] Suppose the server identifies an isolated object (such as a key on the ground), and the pre-defined narrative template is a detective mystery. When generating the story text, the first processing model can output the text "A key was found on the ground, which may be an important clue in the whole incident"; when generating the comic, the second processing model can output a storyboard of "finding the clue" and a close-up of the "key" object; when generating the video, the third processing model can output a video frame of "finding the clue" and give a specific shot of the "key".
[0073] It should be understood that for the same image and / or video to be processed, the generated playback information may have different plot content or style when different preset theme templates are selected.
[0074] For example, the scenario setting is as follows: Scene: Nighttime, a somewhat deserted city street with streetlights.
[0075] Role: Character A: Standing still under a street lamp.
[0076] Character B: Runs quickly towards Character A from a distance on the street.
[0077] Data obtained by the server when calculating the scenario: Proximity: Decreases rapidly.
[0078] Rate of change: Character B has a very high speed value.
[0079] Initial tension score: Based on the elements of "night", "empty streets" and "rapid approach", the server gives a high initial tension score, such as 75 / 100.
[0080] When the preset narrative template is set to Romantic, the server should reduce the weight of the "fast speed" and "night" elements in calculating "danger level"; increase the weight of the "proximity" event "two characters are about to come into contact" and associate it with "emotional intensity". The narrative logic can be set as follows: if the tension score is high, the characters approach quickly, and the theme is "Romantic," the event can be interpreted as "the excitement of a long-awaited reunion," "an impatient confession," or "running into the arms of a lover." In this case, the high tension score (75 / 100) is not considered "dangerous" but is relabeled as an "emotional intensity score." When generating the comic, a soft filter can be used, featuring close-ups of character A's expectant or surprised expression and character B's longing gaze; heart symbols or twinkling stars can be used as background decorations. The comic's line style can be smooth and soft. When generating the video, soothing and melodious orchestral music can be used.
[0081] When the preset narrative template is set to horror, the server should significantly increase the weight of elements such as "fast speed," "night," and "open space" when calculating "danger" and "threat." Character B's rapid movement will be given a very high weight for "aggression" or "threat." A tension score of 75 will be further amplified, potentially triggering a "danger peak" event. The narrative logic can be set as follows: if the tension score is high, the character approaches rapidly, and the theme is "horror," this event can be interpreted as "pursuit," "attack," or "an unknown threat approaching." When generating the comic, specific compositions can be used to create unease. Extensive use of shadows is possible; Character B's face can be deliberately blurred or hidden in shadow. Close-up shots will show Character A's terrified eyes. The comic's line style can be rough, messy, and powerful. When generating the video, discordant, sharp string music, or simply amplified heartbeats and footsteps, can be used.
[0082] It should be understood that in the above approach, the story text, comic, and video may not necessarily be related; that is, the story text, comic, and video can generate different storylines for the same video to be processed. To improve the correlation between the text, comic, and video, the following approach can also be adopted: In one exemplary scheme, the server can call the first processing model and the preset narrative template to generate text based on the real-time narrative analysis framework to obtain story text; then call the fourth processing model to generate comics or videos based on the story text.
[0083] In another exemplary scheme, the server can call the first processing model and the preset narrative template to generate text based on the real-time narrative analysis framework to obtain story text; then call the fourth processing model to generate comics based on the story text; and finally call the fifth processing model to generate animated videos based on the comics.
[0084] It should be understood that the first, second, third, fourth, and fifth processing models mentioned above can be existing natural language processing models or image generation models, and are not specifically limited here. The second processing model can also be a storyboard syntax tree.
[0085] Optionally, after generating the aforementioned playback information, the server can output the playback information in a multimodal manner.
[0086] In one possible implementation, when the information to be played is a story script, multi-character voice synthesis is performed on the story script to obtain a voice script, and the voice script is played through a speaker; When the information to be played is a comic, augmented reality overlay projection is used on the comic. When the information to be played is a video, augmented reality overlay projection is used on the video.
[0087] Optionally, when the information to be played includes a story script, a comic, and a video, the story script can be converted into an audio script and then played simultaneously with the comic and video, thus enabling split-screen playback.
[0088] Furthermore, if there is a connection between the story script, comic, and video, the audio script converted from the story text can be played synchronously as narration for the comic or video.
[0089] Optionally, feedback may also be provided based on the user's interaction behavior while the information to be played is being played. In one exemplary solution, if the above solution is applied to a smart glasses scenario, when the smart glasses are used to play the information to be played, the smart glasses can recognize the user's gestures and / or eye focus, and perform playback operations on the information to be played based on the corresponding gestures and / or eye focus. For example, pause, delete, fast forward, etc.
[0090] Based on the above description, the following will be used as... Figure 3 The flowchart shown in the smart glasses scenario illustrates the data processing flow of this application.
[0091] When a user wears smart glasses and activates the creative creation function, a data stream (which can be either video or a collection of images, including color and depth information) is processed through the smart glasses' camera. The smart glasses or a cloud server then perform real-time semantic segmentation on this data stream to obtain pixel-level label masks (e.g., each pixel-level mask is labeled with a CLOSAT element). Based on these pixel-level label masks, CLOSAT elements are extracted, and structured data is generated. Dynamic narrative decisions (a real-time narrative analysis framework) are constructed based on this structured data. Finally, appropriate playback information is generated by calling the corresponding processing model. When generating a comic, a storyboard syntax tree is used (the specific process could be generating comic panels first, then generating a frame sequence based on the comic panels). The comic is then displayed on the smart glasses' interface using augmented reality technology. During the display, playback can also be initiated via eye-tracking focus coordinates and / or hand gestures to obtain user interaction feedback. When generating story text, emotional tags can be incorporated, and then a voice script can be generated based on multi-character voice synthesis technology, which can then be played back via smart glasses. During presentation, comic playback can be executed through eye-tracking focus coordinates and / or gestures to obtain user interaction feedback. Simultaneously, the smart glasses can adjust the parameters for generating playback information based on user feedback. For example, the weights of various element tags in dynamic narrative decision-making can be dynamically adjusted based on user feedback.
[0092] To better implement the data processing method in the embodiments of this application, based on the data processing method, the embodiments of this application also provide a data processing apparatus, such as... Figure 4 As shown, the data processing device 400 includes: The acquisition module 401 is used to acquire the image and / or video to be processed; The processing module 402 is used to perform semantic segmentation on the image and / or video to be processed, so as to obtain semantic segmentation entities and the element labels corresponding to the semantic segmentation entities. The element labels are used to characterize the narrative elements corresponding to each semantic segmentation entity. The generation module 403 is used to generate playback information based on the semantic segmentation entity and the element label corresponding to the semantic segmentation entity.
[0093] In this embodiment, semantic segmentation is performed on the acquired image and / or video to obtain semantic segmentation entities and corresponding element tags. Finally, creative output information is generated based on the semantic segmentation entities and corresponding element tags, thereby realizing creative output of real-world scenes, providing creative interactive scenarios, and increasing user experience.
[0094] In some embodiments of this application, the generation module 403 is specifically used for: A real-time narrative analysis framework is generated based on the semantic segmentation entity, the element tags corresponding to the semantic segmentation entity, and the preset narrative template. This real-time narrative analysis framework is used to characterize the narrative logic of the information to be played. The target processing model is invoked to process the real-time narrative analysis framework to obtain the information to be played.
[0095] In this embodiment, a real-time narrative analysis framework is generated based on the semantic segmentation entity and its element tags. Finally, the real-time narrative analysis framework is processed based on AI model algorithms and preset narrative templates to generate the information to be played. This can improve the generation efficiency of creative output and make creative output more diverse.
[0096] In some embodiments of this application, the generation module 403 is specifically used for: The first processing model is invoked to generate text for the real-time narrative analysis framework to obtain the story text; And / or, The second processing model is invoked to generate images from the real-time narrative analysis framework to obtain comics; And / or, The third processing model is invoked to generate video from the real-time narrative analysis framework to obtain the video. The story text, the comic, and the video are the information to be played, and the first processing model, the second processing model, and the third processing model are included in the target processing model.
[0097] In this embodiment, different types of processing models are invoked to generate different types of information to be played, thereby making the creative output more diverse and improving the user experience.
[0098] In some embodiments of this application, the processing module 402 is specifically used for: The fourth processing model is invoked to perform semantic segmentation on the image and / or video to be processed, so as to obtain the semantic segmentation entity and the corresponding feature label assigned to the semantic entity. The fourth processing model is a model trained in combination with the feature label.
[0099] In this embodiment, the semantic segmentation model is trained by combining element labels, thereby enabling the direct output of semantically segmented entities carrying mutual labels, thus improving the efficiency of image data processing.
[0100] In some embodiments of this application, the acquisition module 401 is specifically used to acquire the image to be processed and / or the video to be processed based on the camera of a wearable device, the wearable device including but not limited to smart glasses and smart helmets.
[0101] This application provides an application scenario for using smart devices to collect real-time images, thereby enhancing the user's interactive experience.
[0102] In some embodiments of this application, the processing module 402 is further configured to perform multi-role voice synthesis on the story text when the information to be played is story text, so as to obtain a voice script, and play the voice script through a speaker; When the information to be played is a comic, the comic is overlaid and projected using virtual reality. When the information to be played is a video, a virtual reality overlay projection is applied to the video.
[0103] In this embodiment, different playback methods are used to display different playback information, thereby improving the user's interactive experience.
[0104] In some embodiments of this application, the processing module 402 is further configured to perform a playback operation on the information to be played based on gesture actions and / or eye-tracking focus recognition, so as to obtain feedback information; The factual narrative analysis framework was adjusted based on this feedback.
[0105] In this embodiment of the application, user-machine interaction is provided, and user feedback can be applied to the creation process, making the generated playback information more in line with user needs.
[0106] This application also provides a computer device that integrates any of the data processing apparatuses provided in this application, the computer device comprising: One or more processors; Memory; and One or more applications, wherein the applications are stored in memory and configured to be executed by a processor from the steps of the data processing method in any of the embodiments described above.
[0107] This application also provides a computer device that integrates any of the cell handover configurations provided in this application. For example... Figure 5 As shown, it illustrates a structural schematic diagram of the computer device involved in the embodiments of this application, specifically: The computer device may include components such as a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will understand that... Figure 5 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 501 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 502, and by calling data stored in the memory 502, it performs various functions of the computer device and processes data, thereby providing overall monitoring of the computer device. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 501.
[0108] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.
[0109] The computer equipment also includes a power supply 503 that supplies power to the various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 503 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0110] The computer device may also include an input unit 504, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0111] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 501 in the computer device loads the executable files corresponding to the processes of one or more application programs into the memory 502 according to the following instructions, and the processor 501 runs the application programs stored in the memory 502 to realize various functions, as follows: Acquire images and / or videos to be processed; perform semantic segmentation on the images and / or videos to obtain semantic segmentation entities and element tags corresponding to the semantic segmentation entities, wherein the element tags are used to characterize the narrative elements corresponding to each semantic segmentation entity; generate playback information based on the semantic segmentation entities and the element tags corresponding to the semantic segmentation entities.
[0112] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0113] Therefore, embodiments of this application provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a magnetic disk, or an optical disk, etc. A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in any of the data processing methods provided in embodiments of this application. For example, the computer program loaded by the processor can execute the following steps: Acquire images and / or videos to be processed; perform semantic segmentation on the images and / or videos to obtain semantic segmentation entities and element tags corresponding to the semantic segmentation entities, wherein the element tags are used to characterize the narrative elements corresponding to each semantic segmentation entity; generate playback information based on the semantic segmentation entities and the element tags corresponding to the semantic segmentation entities.
[0114] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.
[0115] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.
[0116] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0117] The data processing method, apparatus, computer device, and computer-readable storage medium provided in the embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A data processing method, characterized in that, include: Get the images and / or videos to be processed; Semantic segmentation is performed on the image and / or video to be processed to obtain semantic segmentation entities and element labels corresponding to the semantic segmentation entities. The element labels are used to characterize the narrative elements corresponding to each semantic segmentation entity. Based on the semantic segmentation entity and the corresponding element tags, information to be played is generated.
2. The method according to claim 1, characterized in that, The step of generating playback information based on the semantic segmentation entity and the corresponding element tag of the semantic segmentation entity includes: A real-time narrative analysis framework is generated based on the semantic segmentation entity, the element tags corresponding to the semantic segmentation entity, and the preset narrative template. The real-time narrative analysis framework is used to characterize the narrative logic of the information to be played. The target processing model is invoked to process the real-time narrative analysis framework to obtain the information to be played.
3. The method according to claim 2, characterized in that, The process of calling the target processing model to process the real-time narrative analysis framework to obtain the information to be played includes: The first processing model is invoked to generate text for the real-time narrative analysis framework to obtain the story text; And / or, The second processing model is invoked to generate images from the real-time narrative analysis framework to obtain a comic. And / or, The third processing model is invoked to generate video using the real-time narrative analysis framework to obtain the video. The story text, the comic, and the video are the information to be played, and the first processing model, the second processing model, and the third processing model are included in the target processing model.
4. The method according to any one of claims 1 to 3, characterized in that, Semantic segmentation is performed on the image and / or video to be processed to obtain semantic segmentation entities and corresponding element labels, including: The fourth processing model is invoked to perform semantic segmentation on the image to be processed and / or the video to be processed, so as to obtain the semantic segmentation entity and the element label corresponding to the semantic segmentation entity. The fourth processing model is a semantic segmentation model trained by combining the element label.
5. The method according to any one of claims 1 to 3, characterized in that, Obtaining the images and / or videos to be processed includes: The image to be processed and / or the video to be processed are acquired using a camera on a wearable device, including but not limited to smart glasses and smart helmets.
6. The method according to any one of claims 1 to 3, characterized in that, After generating playback information based on the semantic segmentation entity and the corresponding element tags of the semantic segmentation entity, the method further includes: When the information to be played is story text, multi-character voice synthesis is performed on the story text to obtain a voice script, and the voice script is played. When the information to be played is a comic, the comic is projected using augmented reality overlay. When the information to be played is a video, the video is projected using augmented reality overlay.
7. The method according to any one of claims 1 to 3, characterized in that, After generating playback information based on the semantic segmentation entity and the corresponding element tags of the semantic segmentation entity, the method further includes: Playback of the information to be played is performed based on gestures and / or eye-tracking focus recognition in order to obtain feedback information; The real-time narrative analysis framework is adjusted based on the feedback information, and the real-time narrative analysis framework is used to characterize the narrative logic of the information to be played.
8. A data processing apparatus, characterized in that, include: The acquisition module is used to acquire images and / or videos to be processed; The processing module is used to perform semantic segmentation on the image to be processed and / or the video to be processed, so as to obtain semantic segmentation entities and element labels corresponding to the semantic segmentation entities. The element labels are used to characterize the narrative elements corresponding to each semantic segmentation entity. The generation module is used to generate playback information based on the semantic segmentation entity and the element tags corresponding to the semantic segmentation entity.
9. A computer device, characterized in that, The computer device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the method of any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the method according to any one of claims 1 to 7.