Method, device and system for determining interested content, medium and electronic equipment
By acquiring and recognizing video frames and audio signals that the user's gaze is focused on in the intelligent cockpit system, the system automatically determines the content that the user is interested in, solving the problem of insufficient targeting of content push in existing technologies and achieving precise content push.
Patent Information
- Application Number
- CN202510975882.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-15
- Publication Date
- 2025-10-31
AI Technical Summary
Existing smart cockpit systems struggle to automatically identify content that users are interested in, resulting in insufficiently targeted content recommendations.
By acquiring video frame sequences and audio signal sequences of the target personnel's gaze points in the cockpit, audio and video recognition technologies are used to identify the content of interest to the user, and semantic understanding and fusion are combined to determine the content of interest to the user.
It enables automatic identification of content that users are interested in, improving the effectiveness and relevance of cockpit interaction content and providing precise content delivery.
Smart Images

Figure CN120873231A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to intelligent cockpit technology, and in particular to a method, apparatus, system, medium, and electronic device for determining content of interest. Background Technology
[0002] With the development of smart cockpit technology, the amount of multimedia content that can be played in the cockpit is becoming increasingly rich. However, this abundance of multimedia content makes it difficult to quickly determine the content of interest to the target users in the cockpit, and makes it difficult to provide them with relevant information that matches their interests. Summary of the Invention
[0003] To address the aforementioned technical problems, this disclosure is proposed. Embodiments of this disclosure provide a method, apparatus, system, medium, and electronic device for determining content of interest, enabling automatic identification of a user's content of interest, thereby accurately pushing such content to the user and improving the effectiveness of cockpit interaction content.
[0004] According to a first aspect of the embodiments of this disclosure, a method for determining content of interest is provided, including:
[0005] In response to the fact that the focus of the target person's gaze at the target time point is located on the display screen in the cockpit, the system acquires the video frame sequence corresponding to the display screen within a preset time period including the target time point, and acquires the audio signal sequence corresponding to the display screen within the preset time period.
[0006] The audio signal sequence is identified to obtain a first identification result, and the video frame sequence is identified to obtain a second identification result;
[0007] Based on the first identification result and the second identification result, the content of interest of the target person is determined.
[0008] According to a second aspect of the present disclosure, an apparatus for determining content of interest is provided, comprising:
[0009] The acquisition module is used to respond to the point of focus of the target person's line of sight at the target time point, including the display screen in the cockpit, to acquire the video frame sequence corresponding to the display screen within a preset time period including the target time point, and to acquire the audio signal sequence corresponding to the display screen within the preset time period.
[0010] The recognition module is used to recognize the audio signal sequence to obtain a first recognition result, and to recognize the video frame sequence to obtain a second recognition result;
[0011] The content determination module is used to determine the content of interest to the target person based on the first identification result and the second identification result. This is based on the embodiments provided above in this disclosure.
[0012] According to a third aspect of the present disclosure, a system for determining content of interest is provided, comprising:
[0013] The display screen installed in the cockpit is used to generate and display corresponding images based on video signals;
[0014] The content of interest determination device installed in the cockpit is used to play corresponding audio based on the audio signal;
[0015] The aforementioned content of interest determination device is used to determine the content of interest of the target personnel inside the cockpit.
[0016] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided, the storage medium storing computer program instructions, which, when executed by a processor, are used to implement the above-described method for determining content of interest.
[0017] According to a fifth aspect of the present disclosure, an electronic device is provided, comprising:
[0018] Processor; memory for storing executable instructions of the processor;
[0019] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the above-described method for determining the content of interest.
[0020] Based on the above embodiments of this disclosure, when it is necessary to determine the content of interest, the system can acquire a video frame sequence corresponding to the displayed screen within a preset time period, including the target time point, based on the target person's gaze being focused on the display screen at a target time point. It can also acquire an audio signal sequence corresponding to the displayed screen within the preset time period. A first recognition result is obtained by identifying the audio signal sequence, and a second recognition result is obtained by identifying the video frame sequence. Based on the first and second recognition results, the content of interest for the target person is determined. This technical solution automatically determines the user's content of interest based on their gaze, thereby accurately pushing relevant content to the user and improving the effectiveness of cockpit interaction content.
[0021] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0022] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0023] Figure 1 This is the system architecture diagram to which this disclosure applies.
[0024] Figure 2 This is a flowchart illustrating a method for determining content of interest provided in an exemplary embodiment of this disclosure.
[0025] Figure 3 This is a flowchart illustrating the process of determining the point of interest in a viewpoint in an exemplary embodiment of the present disclosure.
[0026] Figure 4 This is a schematic flowchart illustrating the process of determining the audio signal recognition result in an exemplary embodiment of the present disclosure for determining content of interest.
[0027] Figure 5 This is a schematic flowchart illustrating the process of determining the video signal recognition result in an exemplary embodiment of the present disclosure for determining content of interest.
[0028] Figure 6 This is a flowchart illustrating a method for determining content of interest provided in another exemplary embodiment of this disclosure.
[0029] Figure 7 This is a schematic diagram of the structure of an apparatus for determining content of interest provided in an exemplary embodiment of this disclosure.
[0030] Figure 8 This is a schematic diagram of the structure of a device for determining content of interest provided in another exemplary embodiment of this disclosure.
[0031] Figure 9 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation
[0032] To explain this disclosure, exemplary embodiments of the disclosure will now be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the disclosure, and not all of them. It should be understood that the disclosure is not limited to exemplary embodiments.
[0033] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0034] This disclosure outlines
[0035] In current intelligent cockpit interaction systems, users typically need to actively search for content they are interested in. The system cannot automatically discover the content that users are interested in, nor can it push content that matches the user profile based on user characteristics, resulting in insufficient targeting of content push.
[0036] Exemplary System
[0037] Figure 1 An exemplary system architecture 100 is shown that can be applied to the methods or apparatus for determining content of interest according to embodiments of the present disclosure.
[0038] like Figure 1 As shown, the system architecture 100 may include a terminal device 101, a network 102, a server 103, and an information acquisition device 104. The network 102 serves as the medium for providing a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.
[0039] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as audio players, video players, web browser applications, instant messaging tools, etc.
[0040] Terminal device 101 can be any electronic device capable of determining the content of interest to the user, including but not limited to mobile terminals such as in-vehicle terminals, mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), etc., as well as fixed terminals such as digital TVs, desktop computers, smart home appliances, etc.
[0041] The information acquisition device 104 can be any device used to collect information (including image data and audio data) in the smart cockpit, including but not limited to at least one of the following: camera, microphone, etc.
[0042] Typically, the terminal device 101 is located within a defined space 105, and the information collection device 104 is associated with the space 105. For example, the information collection device 104 can be located within the space 105 to collect various information such as images and sounds from the user, or it can be located outside the space 105 to collect various information such as images and sounds from the surrounding area. The space 105 can be any defined space, such as the interior of a vehicle or a room.
[0043] Server 103 can be a server that provides various services, such as a background audio server that supports audio played on terminal device 101. The background audio server can process the received user-related information and the information played by the terminal device to determine the content of interest to the target personnel within space 105.
[0044] It should be noted that the method for determining content of interest provided in the embodiments of this disclosure can be executed by the server 103 or by the terminal device 101. Accordingly, the device for determining content of interest can be located in the server 103 or in the terminal device 101. The method for determining content of interest provided in the embodiments of this disclosure can also be executed jointly by the terminal device 101 and the server 103. For example, the step of identifying the video frame sequence corresponding to the screen display within a preset time period and the audio signal sequence corresponding to the screen display to obtain a first identification result and a second identification result is executed by the terminal device 101, while the step of determining the content of interest of the target person is executed by the server 103. Accordingly, the modules included in the device for determining content of interest can be respectively located in the terminal device 101 and the server 103.
[0045] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, servers, and information acquisition devices can be included. For example, if the preset audio library is set locally, the above system architecture may exclude networks and servers, including only terminal devices and information acquisition devices.
[0046] Exemplary methods
[0047] Figure 2 This is a schematic flowchart illustrating a method for determining content of interest provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as... Figure 1 On the terminal device 101 or server 103, such as Figure 2 As shown, it includes the following steps:
[0048] Step 201: In response to the target person's gaze being focused on the display screen in the cockpit at the target time point, acquire the video frame sequence corresponding to the display screen within a preset time period, including the target time point, and acquire the audio signal sequence corresponding to the display screen within the preset time period.
[0049] The target personnel can be any person in the cockpit, or a person in a specific position in the cockpit such as the driver's seat or the person facing the display screen, or a person in the cockpit with specific age characteristics such as being under 12 years old, or a person in the cockpit who turns on or adjusts the multimedia player.
[0050] In this embodiment of the disclosure, the target time point can be the moment when the gaze of the target person is captured. Since the operation of capturing the gaze of the target person is real-time, the target time point can be understood as the current moment. The display screen in the cockpit is a screen in the cockpit used to display multimedia content such as videos, games, news, advertisements, weather, and health information.
[0051] In this embodiment of the disclosure, the preset time period is used to indicate a time period including a target time point. The preset time period includes at least one of a first time period and a second time period. The first time period is a time period of a first preset time length preceding and adjacent to the target time point, and the second time period is a time period of a second preset time length following and adjacent to the target time point. For example, a time period within 10 seconds prior to the current time or a time period within 10 seconds backward from the current time. The first and second preset time lengths can also be other time lengths, and this application does not limit the above features.
[0052] The first preset time length and the second preset time length may be the same or different, and this application does not limit this.
[0053] In this embodiment of the disclosure, when playing audio and video resources in the smart cockpit, the audio re-sampling signal of the audio and video resources and the frame-by-frame sampling signal of the screen image can be cached in the electronic device buffer according to a certain time span. When it is determined that the focus of the target person's gaze is on the display screen in the cockpit, the video frame sequence and audio signal sequence within a preset time period can be obtained from the electronic device buffer.
[0054] Step 202: Identify the audio signal sequence to obtain a first identification result, and identify the video frame sequence to obtain a second identification result.
[0055] The first recognition result can be text information, which may include elements and keywords in the text content identified from the audio signal sequence. For example, if the recognized text content is "*** Water Park has an area of 4,000 square meters, including water slides, surfing areas, etc., and can accommodate thousands of people", then the keyword in the text content can be "*** Water Park".
[0056] It should be noted that audio signal sequences can be identified using Automatic Speech Recognition (ASR) algorithms or models to obtain the first recognition result. Before audio recognition, the audio signal sequence can be processed with echo suppression, ambient noise suppression, and automatic gain control to improve the quality of the audio signal.
[0057] The second recognition result can be text information, which may include content elements corresponding to each video frame. The method for recognizing video frame sequences can employ image recognition methods, typically following these steps: preprocessing each video frame, such as denoising, contrast enhancement, and size normalization, to improve image quality; then extracting image elements from the preprocessed image, such as the objects displayed within. During image recognition, semantic understanding of the text in the video frames can also be performed to analyze the content elements within them, such as the title of a playing film or television clip, the name of an advertised product in a commercial, tourist attractions, or people.
[0058] Step 203: Based on the first and second identification results, determine the content of interest to the target personnel.
[0059] In this embodiment of the disclosure, the first identification result and the second identification result can be semantically understood and fused to obtain the content of interest to the target person.
[0060] For example, if the audio in the multimedia video of the water park is determined to introduce information about the water park based on the first identification result, and the video frame in the multimedia video of the water park contains images of the water park based on the second identification result, then it can be considered that the target person's content of interest includes information related to the water park.
[0061] Based on the above embodiments of this disclosure, when it is necessary to determine the content of interest, the system can acquire a video frame sequence corresponding to the displayed screen within a preset time period, including the target time point, based on the target person's gaze being focused on the display screen at a target time point. It can also acquire an audio signal sequence corresponding to the displayed screen within the preset time period. A first recognition result is obtained by identifying the audio signal sequence, and a second recognition result is obtained by identifying the video frame sequence. Based on the first and second recognition results, the content of interest for the target person is determined. This technical solution automatically determines the user's content of interest based on their gaze, thereby accurately pushing relevant content to the user and improving the effectiveness of cockpit interaction content.
[0062] In some optional examples, Figure 3 This is a flowchart illustrating the process of determining the point of interest in a viewpoint in an exemplary embodiment of this disclosure. Figure 3 As shown, this disclosure provides an illustrative example of how to detect the gaze of a target person in real time, including the following steps.
[0063] Step 301: Obtain the position of the display screen; and obtain the eye image of the target person at the target time point, and obtain the eye position of the target person at the target time point.
[0064] The position of the display screen is used to indicate the location of the display screen in the smart cockpit. The information acquisition device 104 can acquire images of one or more spaces in the smart cockpit, and the position of the display screen can be determined in the images of each space based on the target detection method. Then, the spatial position of the display screen in the smart cockpit can be determined according to the position of the information acquisition device and the camera intrinsic parameters.
[0065] The aforementioned object detection methods can be various object detection algorithms or models such as the YOLO detection algorithm and Region-based Convolutional Neural Networks (R-CNN).
[0066] Among them, the eye image of the target person at the target time point can be set in the target space, such as... Figure 1 The information acquisition device 104 includes images captured by a camera. Specifically, the images acquired by the information acquisition device may include facial images of the target person, perform eye key point detection on the facial images, determine the rectangular area containing all eye key points in the image as the eye area of the target person, and extract the image block containing the eye area of the target person from the facial image as the eye image.
[0067] In this embodiment, since the human eye image includes multiple eye key points, the eye key point detection method can accurately obtain the position of multiple eye key points in the face image, and then quickly extract the image block where the target person's eye area is located from the face image as the eye image.
[0068] In this embodiment of the disclosure, a facial image of a target person captured by a camera at a target time point can be acquired, along with the camera position and camera parameters at that time point; based on the facial image, camera position, and camera parameters, the eye position is determined. The camera can be... Figure 1 Information collection device 104 in the middle.
[0069] In other implementations, the target person's seat (center) position (x, y, z) can be determined first, and then a preset value can be added in the direction of the vertical coordinate (z coordinate), such as adding a height of 60 centimeters, to obtain the eye position (x, y, z + 60).
[0070] It is understandable that there may be one or more displays in a smart cockpit, and the position of each display can be obtained.
[0071] Step 302: Using a gaze estimation model, estimate the gaze direction of the target person based on the eye image, wherein the gaze estimation model is trained based on the eye sample image.
[0072] The gaze estimation model can be a model pre-trained based on eye images, capable of estimating the gaze direction from those images. This gaze estimation model can be implemented using a Convolutional Neural Network (CNN) method.
[0073] In this embodiment of the disclosure, a gaze estimation model can be trained using eye sample images labeled with gaze direction.
[0074] Step 303: Based on eye position and gaze direction, determine the target person's gaze focus area.
[0075] The field of vision (field of vision) of a human eye can be determined based on the position of the eye and the direction of the line of sight. For example, the field of vision of a human eye usually refers to the area of about 15 degrees from the center of the line of sight (the area corresponding to the fovea).
[0076] Step 304: In response to the display screen being located within the area of visual attention, determine that the point of visual attention includes the display screen.
[0077] In this embodiment of the disclosure, it can be determined whether the focus of the gaze includes the display screen by determining whether the position of the display screen is within the field of view of the human eye. If the position of the display screen is within the field of view of the human eye, it is determined that the focus of the gaze includes the display screen; if the position of the display screen is not within the field of view of the human eye, it is determined that the focus of the gaze does not include the display screen, and the gaze of the target person can continue to be detected.
[0078] Based on the embodiments of this disclosure, by means of gaze detection, the gaze focus of the target person can be determined, thereby determining whether the target person's focus is on the display screen in the cockpit. When the target person's focus is on the display screen in the cockpit, the content of interest of the user can be determined by the display screen's display image within a preset time period and the audio signal corresponding to the display image.
[0079] Figure 4 This is a schematic flowchart illustrating the process of determining the audio signal recognition result in an exemplary embodiment of the present disclosure for determining content of interest. Figure 4 As shown, this embodiment uses the method of determining the audio signal recognition result as an example to illustrate the following steps.
[0080] Step 401: Perform audio recognition on the audio signal sequence to obtain the text content corresponding to the audio signal sequence.
[0081] The audio signal sequence includes audio signals within a preset time period. Audio signals in the audio signal sequence can be identified in chronological order to obtain a text content.
[0082] Step 402: Perform natural language processing on the text content and determine the first element of interest based on the results of natural language processing.
[0083] The process involves data cleaning of the text content corresponding to the audio signal sequence, segmenting the cleaned text into meaningful units such as words, phrases, or sentences, and then removing common words that are not meaningful in the text analysis to obtain the remaining effective words.
[0084] The first element of interest (FII) is used to indicate key elements identified from the text content, and may include at least one of information such as time, scene, location, and topic. The FII can be obtained by performing natural language processing and semantic analysis on the text content.
[0085] In this embodiment of the disclosure, the domain and intent of the text content can first be determined based on semantic analysis and semantic understanding, such as what film or television clip is being played, what product advertisement is being shown, or what tourist attraction is being introduced. After determining the domain and intent corresponding to the text content, different key elements can be further extracted from the text content as first interest elements according to different domains.
[0086] For example, if the domain of the text content is determined to be an advertising product based on semantic analysis and semantic understanding, the product name and product function of the advertising product can be obtained from the text content; if the domain of the text content is determined to be a film clip based on semantic analysis and semantic understanding, the film lines and film characters can be obtained from the text content; if the domain of the text content is determined to be a tourist attraction based on semantic analysis and semantic understanding, the attraction name and location can be obtained from the text content.
[0087] Step 403: Determine the first recognition result based on the first interest element.
[0088] In this embodiment of the disclosure, the first interest element can be determined as the first identification result, or the first event can be determined based on the first interest element and the first event can be determined as the first identification result.
[0089] For example, if the first interest element includes "*** Water Park" and "Opening date 2025 / 06 / 27", then a first event "Promotional activities for the opening of *** Water Park on 2025 / 06 / 27" can be generated, and this first event can be identified as the first identification result.
[0090] In this embodiment of the disclosure, a first event can be generated by semantic association of a first interest element.
[0091] Based on the embodiments of this disclosure, by means of audio recognition and natural language processing, a first recognition result including a first interest element can be identified from an audio signal sequence, thereby obtaining more comprehensive content related to the content played on the display screen, which helps to accurately determine the user's interest content based on the first recognition result.
[0092] Figure 5 This is a schematic flowchart illustrating the process of determining the video signal recognition result in an exemplary embodiment of the present disclosure for determining content of interest.
[0093] Step 501: Perform image recognition on each video frame in the video frame sequence to obtain the image recognition result for each video frame, wherein the image recognition result includes the content elements of the corresponding video frame.
[0094] Among them, each video frame in the cached video frame sequence can be identified. Commonly used methods include image segmentation, object detection, person recognition, and text recognition.
[0095] The content elements of a video frame can include information such as the video's background, scene, plot, characters, theme, location, and environment.
[0096] In this embodiment of the disclosure, visual features can be extracted and analyzed from video frames to determine the domain of the video, and different content elements can be obtained according to the different video domains.
[0097] For example, if the video is determined to be an introductory video for a tourist attraction based on the visual characteristics of the video frame, then the video background information and text description information can be determined as the content elements of the video frame; if the video content is determined to be a film clip based on the visual characteristics of the video frame, then the film characters, the actors playing the film characters, and the storyline in the video frame can be determined as the content elements of the video frame.
[0098] In this embodiment of the disclosure, when performing image recognition, the video background features, people, objects, and video text (used to supplement the screen information, such as time, location, and key content such as character introductions) in the video frame can be identified first through image recognition methods. Then, the video background features, video text, video people, and objects are further analyzed and identified to obtain the content elements of the video frame.
[0099] Step 502: Determine the second element of interest based on the content elements of each video frame in the video frame sequence.
[0100] In this embodiment of the disclosure, the domain and intent of the video can be determined first based on the content elements of the video frame, and then a second element of interest can be determined from the content elements of the video frame. For example, if the video is determined to be an advertising product based on the content elements of the video frame, the product name and product function of the advertising product can be obtained from the content elements; if the video is determined to be a film clip based on the content elements of the video frame, the film dialogue and film characters can be obtained from the content elements of the video frame; if the video is determined to be an introduction and promotional video for a tourist attraction based on the content elements of the video frame, the attraction name, attraction location, etc., can be obtained from the content elements of the video frame.
[0101] Step 503: Determine the second recognition result based on the second interest element.
[0102] In this embodiment of the disclosure, the second interest element can be determined as the second recognition result, or the second event can be determined based on the second interest element and then determined as the second recognition result.
[0103] For example, if the second element of interest includes "*** Water Park", "** Valley", and "Opening Date 2025 / 06 / 27", then a second event "Promotional activities for the opening of *** Water Park located in ** Valley on 2025 / 06 / 27" can be generated, and this second event can be identified as the second recognition result. Alternatively, if the second element of interest includes "** Starring", "Journey to the West", and "The Monkey King 2" clip, then a second event "The Monkey King 2" movie introduction starring ** can be generated, and this second event can be identified as the second recognition result.
[0104] Based on the embodiments of this disclosure, image recognition can identify a second recognition result including a second interest element from a video frame sequence, which can provide a more accurate description of the playback object and help to more comprehensively determine the user's content of interest based on the second recognition result.
[0105] In an optional example, after determining the first identification result and the second identification result, a first interest element can be extracted from the first identification result, and a second interest element corresponding to the first interest element in the time dimension can be extracted from the second identification result; then the first interest element and the second interest element are fused according to the time dimension to obtain the content of interest of the target person.
[0106] In this implementation, when determining the target person's content of interest by combining the first and second recognition results, the first and second elements of interest can be fused along the time dimension to ensure that the first and second elements of interest are fused synchronously in time. Typically, recognizing an audio signal sequence within a preset time period may yield multiple first elements of interest with a temporal relationship, while recognizing a video frame sequence within a preset time period may yield multiple second elements of interest with a temporal relationship. To more accurately fuse the first and second elements of interest, they can be fused along the time dimension.
[0107] For example, if two advertisements are played within a preset time period (e.g., 10 seconds or 20 seconds), when identifying the audio signal sequence within the preset time period, a first set of first interest elements related to the first advertisement content is obtained first, and then a set of first interest elements related to the second advertisement content is obtained. Similarly, when identifying the video frame sequence within the preset time period, a second set of second interest elements related to the first advertisement content is obtained first, and then a set of second interest elements related to the second advertisement content is obtained. When fusing according to the time dimension, the first interest elements related to the first advertisement content and the second interest elements related to the first advertisement content can be fused, and the first interest elements related to the second advertisement content and the second interest elements related to the second advertisement content can be fused.
[0108] In order to fuse the first and second elements of interest according to the time dimension, the audio time points corresponding to each first element of interest and the video time points (video frames) corresponding to each second element of interest can be stored separately, and then the first and second elements of interest can be fused according to the time points.
[0109] In an optional example, the first interest element includes a first event element, and the second interest element includes a second event element. When fusing the first interest element and the second interest element according to the time dimension to obtain the content of interest to the target person, the first event element and the second event element can be fused into event elements. Based on the event element fusion result, the events of interest to the target person can be determined. Based on the events of interest, the content of interest to the target person can be determined.
[0110] The first interest element can include at least one of various types of elements, such as time, place, people, cause, process, and result. The second interest element can also include at least one of various types of elements, such as time, place, people, cause, process, and result. When the first and second interest elements are merged, the complete elements of an event of interest can be obtained by merging the event elements, and the content of interest of the target personnel can be determined based on the complete elements of the event of interest.
[0111] In an optional example, based on events of interest, the content of interest of the target person is determined, including: obtaining a profile of the target person; and based on the profile and events of interest, determining the content of interest of the target person.
[0112] The target person profile can include the target person's basic attributes (age, gender, region, education background, occupation, income, marital status), consumption behavior (purchase history, shopping frequency, consumption amount, product features they care about such as price sensitivity, functional needs, etc.), social media preferences (social platform usage habits, interactive behaviors (such as likes, comments, shares), areas of interest (such as technology, culture, maternal and infant care, etc.), and information such as community affiliation.
[0113] The technical solution disclosed herein involves the collection, storage, use, processing, transmission, provision, and disclosure of users' personal information, including basic attributes, consumption behavior, and social media preferences, all of which comply with relevant laws and regulations and do not violate public order and good morals. Furthermore, the collection and use of users' personal information in this technical solution are conducted with the user's knowledge and authorization, and do not involve the illegal collection or use of users' personal information.
[0114] Based on the person's profile and the events of interest, we can determine why the target person is interested in the events and which aspects of the events they are more concerned about. From this, we can obtain the aspects of the events that the target person might be interested in from the web server as the content of interest for the target person.
[0115] For example, if the event of interest is a promotional event for a tourist attraction, and the target audience is determined to be a self-driving travel enthusiast based on the target audience's profile, then all aspects that the tourist enthusiast might be interested in related to the tourist attraction can be identified as the target audience's content of interest, such as basic expenses such as transportation, accommodation, and dining near the attraction, as well as additional costs such as attraction tickets and special experiences, the tourist flow, local characteristics, and cultural customs of the attraction, etc.
[0116] In this implementation, by combining a person's profile and events of interest, more concrete and targeted information related to the events of interest and matching the person's profile can be recommended as content of interest to the target person.
[0117] In one optional example, the person's profile can be updated based on the identified content of interest.
[0118] In this implementation, after determining the target personnel's interests, the person profile is updated based on those interests. This helps to update the target personnel's latest data in real time, adjust the person profile in a timely manner, and ensure that the determined interests always remain valid.
[0119] Figure 6 This is a flowchart illustrating a method for determining content of interest provided in another exemplary embodiment of this disclosure.
[0120] Step 601: In response to the target person's gaze being focused on the display screen in the cockpit at the target time point, acquire the video frame sequence corresponding to the display screen within a preset time period, including the target time point, and acquire the audio signal sequence corresponding to the display screen within the preset time period.
[0121] Step 602: Identify the audio signal sequence to obtain a first identification result, and identify the video frame sequence to obtain a second identification result.
[0122] Step 603: Based on the first and second identification results, determine the content of interest to the target personnel.
[0123] In this embodiment of the disclosure, the implementation methods of steps 601 to 603 can be found in [reference needed]. Figure 2 The description of the illustrated embodiments will not be repeated here.
[0124] Step 604: Obtain related information that matches the content of interest.
[0125] In this embodiment of the disclosure, relevant information matching the content of interest can be obtained through a web server or other social media platforms. For example, if it is determined that the target person's content of interest is the basic expenses such as transportation, accommodation, and catering near the scenic spot, as well as additional costs such as scenic spot tickets and special experiences, and information such as the scenic spot's visitor flow, local characteristics, and cultural customs, then the above information can be obtained from the scenic spot's official web server or other social media platforms as relevant information.
[0126] The associated information can be content in various media formats such as images, videos, and audio.
[0127] Step 605: Display the associated information on the screen.
[0128] Based on the embodiments of this disclosure, after automatically determining the content of interest to the user according to the user's gaze, the content of interest to the user can be automatically retrieved, and services or consultations can be proactively provided to the user, which greatly expands the connotation of cockpit interaction.
[0129] Any of the methods for determining content of interest provided in this disclosure can be executed by any suitable device with data processing capabilities, including but not limited to: terminal devices and servers. Alternatively, any of the methods for determining content of interest provided in this disclosure can be executed by a processor, such as by a processor executing any of the methods for determining content of interest mentioned in this disclosure by calling corresponding instructions stored in memory. Further details will not be elaborated below.
[0130] Exemplary device
[0131] Figure 7 This is a schematic diagram of the structure of an apparatus for determining content of interest provided in an exemplary embodiment of this disclosure. Figure 7 As shown, the device for determining content of interest may include:
[0132] The acquisition module 71 is used to respond to the focus of the target person's line of sight at the target time point, including the display screen in the cockpit, to acquire the video frame sequence corresponding to the display screen within a preset time period including the target time point, and to acquire the audio signal sequence corresponding to the display screen within the preset time period.
[0133] The recognition module 72 is used to recognize the audio signal sequence to obtain a first recognition result, and to recognize the video frame sequence to obtain a second recognition result;
[0134] The content determination module 73 is used to determine the content of interest to the target personnel based on the first identification result and the second identification result.
[0135] The acquisition module 71 mentioned above may include a processing module and an image acquisition device. The image acquisition device can be any type of image sensor, such as a camera or video camera, as long as it can acquire image and video data. The processing module can determine whether the target person's gaze at the target time point includes the display screen inside the cockpit based on the image data acquired by the image acquisition device. If it determines that the target person's gaze at the target time point includes the display screen inside the cockpit, it reads the video frame sequence and audio signal sequence from the buffer.
[0136] The aforementioned recognition module 72 may include a module for recognizing audio signals and a module for recognizing video signals. The module for recognizing audio signals can use an audio signal recognition model to recognize the audio signal sequence and obtain a first recognition result. The module for recognizing video signals can use an image recognition algorithm to perform image recognition on each video frame and obtain a second recognition result.
[0137] In an optional example, the module used to identify audio signals can be a functional module within the processor. The processing module can call an audio signal recognition model or audio signal recognition algorithm stored in memory to identify the audio signal sequence.
[0138] In an optional example, the module for recognizing the video signal can be a processing module within a processor. This processing module can call image recognition algorithms or models stored in memory to recognize the audio signal sequence. The content determination module 73 can be a processing module capable of analyzing, fusing, and semantically associating the first and second recognition results to determine the content of interest to the target user.
[0139] The processing module can be a functional module in the processor or a separate processing chip.
[0140] The electrical connection between the processor and the memory can be at least one of the following: a data bus connection, a control signal line connection, and an address bus connection.
[0141] Figure 8 This is a schematic diagram of the structure of a device for determining content of interest provided in another exemplary embodiment of this disclosure. In the above... Figure 7 Based on the illustrated embodiment, in some implementations, the acquisition module 71 may include:
[0142] The acquisition unit 711 is used to acquire an eye image of the target person at a target time point, acquire the eye position of the target person at the target time point, and acquire the position of the display screen;
[0143] The gaze estimation unit 712 is used to estimate the gaze of the target person by using the gaze estimation model to obtain the gaze direction of the target person. The gaze estimation model is trained based on the eye sample images.
[0144] The gaze attention area determination unit 713 is used to determine the gaze attention area of a target person based on eye position and gaze direction;
[0145] The attention point determination unit 714 is used to determine that the attention point includes the display screen in response to the position of the display screen being within the field of view.
[0146] The acquisition unit 711 can access the target detection model or target detection algorithm in the memory, extract the eye image from the spatial image, and obtain the eye position of the target person at the target time point through back projection based on the eye image and the position of the image acquisition device that acquired the eye image, as well as the camera's internal and external parameters. The acquisition unit 711 can be an independent processing chip or a functional module (processing unit / processing module) within a processing chip.
[0147] The gaze estimation unit 712, the gaze attention area determination unit 713, and the attention point determination unit 714 can all be independent processing chips, or they can be different processing units in a single processing chip.
[0148] In some implementations, the acquisition unit 711 is specifically used to acquire a facial image of the target person captured by the camera at a target time point, and to acquire the camera position and camera parameters at the target time point; and to determine the eye position based on the facial image, camera position and camera parameters.
[0149] In some implementations, the preset time period includes at least one of a first time period and a second time period, wherein the first time period is a first preset time length period preceding and adjacent to the target time point, and the second time period is a second preset time length period following and adjacent to the target time point.
[0150] In some implementations, the identification module 72 may include:
[0151] The audio recognition unit 721 is used to perform audio recognition on the audio signal sequence to obtain the text content corresponding to the audio signal sequence.
[0152] Processing unit 722 is used to perform natural language processing on text content and determine the first interest element based on the result of natural language processing;
[0153] The first determining unit 723 is used to determine the first identification result based on the first interest element.
[0154] In some implementations, the identification module 72 may include:
[0155] The image recognition unit 724 is used to perform image recognition on each video frame in the video frame sequence to obtain the image recognition result of each video frame, wherein the image recognition result includes the content elements of the corresponding video frame;
[0156] The second determining unit 725 is used to determine the second interest element based on the content elements of each video frame in the video frame sequence;
[0157] The third determining unit 726 is used to determine the second recognition result based on the second interest element.
[0158] The identification module 72, the second determination unit 725, and the third determination unit 726 can all be independent processing chips or different processing units within a single processing chip.
[0159] In some implementations, the content determination module 73 may include:
[0160] Extraction unit 731 is used to extract a first interest element from the first recognition result and extract a second interest element that corresponds to the first interest element in the time dimension from the second recognition result.
[0161] The fusion unit 732 is used to fuse the first interest element and the second interest element according to the time dimension to obtain the content that the target person is interested in.
[0162] In some implementations, the first element of interest includes a first event element, and the second element of interest includes a second event element;
[0163] The fusion unit 732 can be used to fuse the first event element and the second event element, determine the events of interest of the target personnel based on the event element fusion result, and determine the content of interest of the target personnel based on the events of interest.
[0164] In some implementations, the fusion unit 732 can be used to acquire a profile of the target person; based on the profile and events of interest, to determine the content of interest of the target person.
[0165] In some embodiments, the content of interest determination device may further include: a portrait update module 74, used to update the portrait of a person based on the determined content of interest.
[0166] The fusion unit 732 can be an independent processing chip or a processing unit within a processing chip.
[0167] In some embodiments, the apparatus for determining the content of interest may further include:
[0168] The associated information acquisition module 75 is used to acquire associated information that matches the content of interest;
[0169] Display module 76 is used to display related information via a display screen.
[0170] The associated information acquisition module 75 can access a network server via the network to obtain associated information matching the content of interest. The display module 76 can include multiple output modules and a display screen, and each output module can send the associated information to the display screen for information display via a data channel.
[0171] It should be noted that the modules and units in this device can be disassembled and / or recombined, and these disassemblies and / or recombinations should be considered as equivalent solutions of this device.
[0172] It should be noted that the specific implementation of the device for determining the content of interest in this disclosure is similar to the specific implementation of the method for determining the content of interest in this disclosure. For details, please refer to the section on determining the content of interest. To reduce redundancy, further details will not be provided.
[0173] Exemplary electronic devices
[0174] Figure 9 A structural diagram of an electronic device provided in an embodiment of this disclosure includes at least one processor 11 and a memory 12.
[0175] The processor 11 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 9 to perform desired functions.
[0176] The memory 12 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 11 may execute one or more computer program instructions to implement the methods and / or other desired functions determined in the various embodiments of this disclosure above.
[0177] In one example, the electronic device 9 may also include an input device 13 and an output device 14, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0178] The input device 13 may also include, for example, a keyboard, a mouse, etc.
[0179] The output device 14 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0180] Of course, for the sake of simplicity, Figure 9 Only some of the components of the electronic device 9 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 9 may include any other suitable components depending on the specific application.
[0181] Exemplary systems, computer program products, and computer-readable storage media
[0182] In addition to the methods and devices described above, embodiments of this disclosure may also provide a content of interest determination system, including: a display screen disposed in the cockpit for generating and displaying corresponding images based on video signals; and a content of interest determination device disposed in the cockpit for playing corresponding audio based on audio signals; the above-mentioned Figure 7 and Figure 8 The device for determining the content of interest is used to determine the content of interest of a target person inside the cockpit.
[0183] Embodiments of this disclosure may also provide a computer program product, including computer program instructions that, when executed by a processor, cause the processor to perform the steps of the content of interest determination method described in the various embodiments of this disclosure in the "Exemplary Methods" section above.
[0184] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. These programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0185] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions that, when executed by a processor, cause the processor to perform steps in the methods for identifying targets in an image according to the various embodiments of this disclosure described in the "Exemplary Methods" section above.
[0186] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may include, but is not limited to, systems, apparatuses, or devices that are electrical, magnetic, optical, electromagnetic, infrared, or semiconductor, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0187] The basic principles of this disclosure have been described above with reference to specific embodiments. However, the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0188] Those skilled in the art can make various modifications and variations to this disclosure without departing from the spirit and scope of this application. Therefore, this disclosure is also intended to include such modifications and variations if they fall within the scope of the claims of this disclosure and their equivalents.
Claims
1. A method for determining content of interest, comprising: In response to the fact that the focus of the target person's gaze at the target time point is located on the display screen in the cockpit, the system acquires the video frame sequence corresponding to the display screen within a preset time period including the target time point, and acquires the audio signal sequence corresponding to the display screen within the preset time period. The audio signal sequence is identified to obtain a first identification result, and the video frame sequence is identified to obtain a second identification result; Based on the first identification result and the second identification result, the content of interest of the target person is determined.
2. The method according to claim 1, wherein, Before acquiring the video frame sequence corresponding to the displayed screen within the cockpit at the target time point when the target person's line of sight is at the target time point, and acquiring the audio signal sequence corresponding to the displayed screen within the preset time period, the method further includes: The position of the display screen is obtained; and the eye image of the target person at the target time point is obtained, and the eye position of the target person at the target time point is obtained; Using a gaze estimation model, the gaze direction of the target person is estimated based on the eye image, wherein the gaze estimation model is trained based on eye sample images; Based on the eye position and the direction of gaze, determine the area of focus of the target person's gaze; In response to the fact that the position of the display screen is within the area of visual attention, it is determined that the point of visual attention includes the display screen.
3. The method according to claim 2, wherein, The step of obtaining the eye position of the target person at the target time point includes: The camera captures a facial image of the target person at the target time point, and the camera position and camera parameters at the target time point are also acquired. The eye position is determined based on the facial image, the camera position, and the camera parameters.
4. The method according to any one of claims 1-3, wherein, The preset time period includes at least one of a first time period and a second time period. The first time period is a time period of a first preset time length that is before and adjacent to the target time point, and the second time period is a time period of a second preset time length that is after and adjacent to the target time point.
5. The method according to claim 1, wherein, The step of identifying the audio signal sequence to obtain a first identification result includes: Perform audio recognition on the audio signal sequence to obtain the text content corresponding to the audio signal sequence; Natural language processing is performed on the text content, and a first interest element is determined based on the result of the natural language processing; Based on the first interest element, the first recognition result is determined.
6. The method according to claim 5, wherein, The step of identifying the video frame sequence to obtain a second identification result includes: Image recognition is performed on each video frame in the video frame sequence to obtain an image recognition result for each video frame, wherein the image recognition result includes the content elements of the corresponding video frame; Based on the content elements of each video frame in the video frame sequence, a second interest element is determined; The second recognition result is determined based on the second interest element.
7. The method according to claim 6, wherein, The step of determining the target person's interests based on the first identification result and the second identification result includes: Extract the first interest element from the first identification result, and extract the second interest element corresponding to the first interest element in the time dimension from the second identification result; The first interest element and the second interest element are fused according to the time dimension to obtain the content of interest to the target person.
8. The method according to claim 7, wherein, The first interest element includes a first event element, and the second interest element includes a second event element; The process of fusing the first interest element and the second interest element along a time dimension to obtain the content of interest to the target person includes: The first event element and the second event element are fused together, and the events of interest to the target person are determined based on the event element fusion result. Based on the events of interest, determine the content of interest to the target personnel.
9. The method according to claim 8, wherein, The step of determining the target person's interests based on the events of interest includes: Obtain a portrait of the target person; Based on the person profile and the events of interest, the content of interest to the target person is determined.
10. The method according to claim 8 or 9, wherein, After determining the target person's interests based on the person's profile and the events of interest, the method further includes: The portrait of the person is updated based on the identified content of interest.
11. The method according to any one of claims 1-10, wherein, After determining the target person's interests based on the first and second identification results, the method further includes: Obtain association information that matches the content of interest; The associated information is displayed on the screen.
12. An apparatus for determining content of interest, comprising: The acquisition module is used to respond to the point of focus of the target person's line of sight at the target time point, including the display screen in the cockpit, to acquire the video frame sequence corresponding to the display screen within a preset time period including the target time point, and to acquire the audio signal sequence corresponding to the display screen within the preset time period. The recognition module is used to recognize the audio signal sequence to obtain a first recognition result, and to recognize the video frame sequence to obtain a second recognition result; The content determination module is used to determine the content of interest to the target person based on the first identification result and the second identification result.
13. The apparatus according to claim 12, wherein, The acquisition module includes: The acquisition unit is used to acquire an eye image of the target person at the target time point, acquire the eye position of the target person at the target time point, and acquire the position of the display screen; A gaze estimation unit is used to estimate the gaze of the target person by using a gaze estimation model to obtain the gaze direction of the target person. The gaze estimation model is trained based on eye sample images. A gaze attention area determination unit is used to determine the gaze attention area of the target person based on the eye position and the gaze direction. The attention point determination unit is configured to determine that the attention point includes the display screen in response to the position of the display screen being within the visual attention area.
14. A system for determining content of interest, comprising: The display screen installed in the cockpit is used to generate and display corresponding images based on video signals; The content of interest determination device installed in the cockpit is used to play corresponding audio based on the audio signal; The apparatus for determining content of interest according to any one of claims 12-13 is used to determine the content of interest of a target person in the cockpit.
15. A computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the method for determining content of interest as described in any one of claims 1-11.
16. An electronic device comprising: processor; Memory for storing the executable instructions of the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the content of interest determination method according to any one of claims 1-11.