Cboth case generation method and device, electronic equipment and storage medium
By identifying objects and scenes in user images and combining them with sentiment analysis to generate copy, the problem of long copy creation time is solved and users' social activity is improved.
Patent Information
- Application Number
- CN202510793947.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-16
AI Technical Summary
In the existing technology, copywriting takes a long time, resulting in low social activity of users.
By identifying objects and scenes in the user's target image, obtaining object labels and scene labels, performing sentiment analysis, and combining user information to generate target copy.
It enables the rapid generation of copy that meets the emotional needs of users and improves their social activity.
Smart Images

Figure CN120654667A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence processing technology, and in particular to a copywriting generation method, device, electronic device and storage medium. Background Art
[0002] Copywriting encompasses all textual communication in brand marketing. It includes, but is not limited to, advertising slogans, promotional messages, social media content, product descriptions, marketing emails, social media posts, and website brochures. Copywriting is a blend of art and strategy, requiring a deep understanding of the target audience, a clear brand message, and both creativity and expressiveness. Social copywriting refers to short, concise textual content published on social media platforms to convey information, express opinions, and promote products. It often incorporates multimedia elements such as images, videos, and animations to capture the audience's attention.
[0003] In real-world scenarios, whether it is the creation of copywriting or social copywriting, users need to spend a lot of time. However, many users do not want to spend a lot of time, which leads to low activity on their social software.
[0004] Therefore, the existing technology has the problem that users' social activity is low due to the long time consumed in creating copywriting. Summary of the Invention
[0005] The embodiments of the present invention provide a copy generation method, device, electronic device and storage medium, which are intended to automatically generate copy based on images.
[0006] To solve the above technical problems, the embodiments of the present invention provide the following technical solutions: A copywriting generation method, comprising: Get the user's target image and user information; Performing object recognition and scene recognition on the target image respectively to obtain an object label and a scene label of the target image; Performing sentiment analysis on the object label and the scene label to obtain a sentiment analysis result, and correcting the sentiment analysis result according to the user information to obtain a sentiment label for the target image; A target text for the target image is generated according to the emotion label, the object label, and the scene label.
[0007] Optionally, the user information includes a preference database of the user; performing sentiment analysis on the object label and the scene label to obtain a sentiment analysis result, and correcting the sentiment analysis result according to the user information to obtain the sentiment label of the target image includes: Get the upload time of the target image; Based on the correspondence between the preset labels and the emotions, obtaining the object emotion label of the object label and the scene emotion label of the scene label; determining a time emotion tag of the user at the uploading time according to the preference database; Performing emotion value frequency statistics on the object emotion label, the scene emotion label, and the time emotion label to obtain emotion value statistics of the target image; Determine the emotion label of the target image according to the emotion value statistics.
[0008] Optionally, performing emotion value frequency statistics on the object emotion label, the scene emotion label, and the time emotion label to obtain the emotion value statistics result of the target image includes: Performing emotion value frequency statistics on the object emotion label, the scene emotion label, and the time emotion label to obtain an initial emotion value statistical result of the target image; Obtaining the high-frequency emotion tag that appears most frequently in the preference database, and the high-frequency emotion frequency at which the high-frequency emotion tag appears; Determining a correction frequency based on a preset correction weight and the high-frequency emotion frequency; The one or more emotions that appear most frequently in the initial emotion value statistics are corrected according to the correction frequency to obtain the emotion value statistics of the target image.
[0009] Optionally, the user information includes a target release time of the target copy and a preference database of the user; and the step of correcting the result of sentiment analysis according to the user information to obtain the sentiment label of the target image further includes: Obtaining a time emotion label of the target release time in the preference database; Determine the time emotion tag as the emotion tag of the user.
[0010] Optionally, the target publishing time also includes the time period in which the target publishing time falls; and the step of obtaining the time emotion tag of the target publishing time in the preference database further includes: Determining whether the target publishing time is within a preset birthday period of the user; If so, positive emotion is selected as the temporal emotion label; If not, determining whether the target release time is during a pre-set holiday; When the target release time is the holiday, determining the time emotion tag corresponding to the holiday according to the preset correspondence between holidays and emotion tags; When the target publishing time is not the holiday, determining the time emotion label according to the time period of the target publishing time; The time period is obtained by dividing a day into multiple time intervals, and the time emotion labels corresponding to each time period are not completely consistent.
[0011] Optionally, at least one of the emotion label, the object label, and the scene label includes multiple labels; and generating a target text for the target image based on the emotion label, the object label, and the scene label includes: Merging the emotion label, the object label, and the scene label into a label set; Randomly selecting one of the object labels, one of the scene labels, and one of the emotion labels from the label set and combining them to obtain multiple text label sets; Generating multiple texts for the target image according to the multiple text tag sets; A target copy for the target image is selected from the multiple copies.
[0012] Optionally, the user information includes spoken language data of the user; after generating a target text for the target image according to the emotion tag, the object tag, and the scene tag, the method further includes: Performing colloquial conversion on the target text according to the colloquial data to obtain a colloquial text; The colloquial data includes at least one of dialects, catchphrases, internet slang, slang and idiomatic expressions.
[0013] A copywriting generating device, comprising: Data acquisition module, used to obtain the user's target image and user information; a label recognition module, configured to perform object recognition and scene recognition on the target image, respectively, to obtain an object label and a scene label of the target image; A sentiment analysis module is used to perform sentiment analysis on the object label and the scene label to obtain a sentiment analysis result, and to modify the sentiment analysis result according to the user information to obtain a sentiment label for the target image; A text generation module is used to generate a target text for the target image based on the emotion label, the object label and the scene label.
[0014] An electronic device includes a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps: Get the user's target image and user information; Performing object recognition and scene recognition on the target image respectively to obtain an object label and a scene label of the target image; Performing sentiment analysis on the object label and the scene label to obtain a sentiment analysis result, and correcting the sentiment analysis result according to the user information to obtain a sentiment label for the target image; A target text for the target image is generated according to the emotion label, the object label, and the scene label.
[0015] A computer-readable storage medium stores a computer program, which is loaded by a processor to execute the steps in the above-mentioned copywriting generation method.
[0016] In an embodiment of the present invention, the object and scene in the target picture of the user are identified to obtain the object label and scene label of the target picture; the emotion label of the target picture is obtained by performing sentiment analysis on the object label and scene label, and correcting the result of the sentiment analysis according to the user information, so as to automatically identify the emotional tendency of the user in sending the target picture; the target copy of the target picture is generated according to the emotion label, object label and scene label, thereby realizing the automatic generation of the target copy that meets the emotional needs of the user according to the picture; since the time of obtaining the target picture of the user is very short and the target copy has a high degree of consistency with the user's emotions, it can not only better meet the user's needs, but also quickly obtain the target copy that meets the requirements, thereby increasing the user's enthusiasm for sending social copy and improving his social activity. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 A schematic diagram of a scenario of an embodiment of a copywriting generation system provided by an embodiment of the present invention; Figure 2 A schematic diagram of another embodiment of the copywriting generation system provided by an embodiment of the present invention; Figure 3 A schematic diagram of a flow chart of an embodiment of a method for generating a document according to an embodiment of the present invention; Figure 4 A flowchart illustrating an embodiment of the present invention providing a social software that automatically generates text based on a text generation method; Figure 5 A schematic structural diagram of an embodiment of a document generation device provided by an embodiment of the present invention; Figure 6 It is a structural diagram of an embodiment of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts shall fall within the scope of protection of the present invention.
[0020] In the following description, the specific embodiments of the present invention will be described with reference to steps and symbols performed by one or more computers, unless otherwise specified. Therefore, these steps and operations will be mentioned several times as being performed by a computer, and the computer execution referred to herein includes the operation of a computer processing unit by electronic signals representing data in a structured form. This operation converts the data or maintains it at a location in the computer's memory system, which can be reconfigured or otherwise change the operation of the computer in a manner familiar to testers in the field. The data structure in which the data is maintained is a physical location in the memory, which has specific characteristics defined by the data format. However, the principles of the present invention are described in the above text, which does not represent a limitation, and testers in the field will understand that the various steps and operations described below can also be implemented in hardware.
[0021] As used herein, the terms "module" or "unit" may be considered software objects executed on the computing system. The various components, modules, engines, and services described herein may be considered implementation objects on the computing system. While the devices and methods described herein are preferably implemented in software, they may also be implemented in hardware and fall within the scope of protection of the present invention.
[0022] Embodiments of the present invention provide a document generation method, device, electronic device, and storage medium.
[0023] See also Figure 1 , Figure 1 The present invention provides a scenario diagram of an embodiment of a document generation system. The document generation system may include a client 100 and a server 200. The client 100 and the server 200 are connected via a network. The server 200 is integrated with a document generation device. The server 200 may be a work platform server (i.e., a server loaded with a work platform). Figure 1The client 100 can access the server 200. In the embodiment of the present invention, the server 200 is mainly used to obtain the user's target image and user information; perform object recognition and scene recognition on the target image to obtain the object label and scene label of the target image; perform sentiment analysis on the object label and scene label to obtain a sentiment analysis result, and modify the sentiment analysis result according to the user information to obtain the sentiment label of the target image; and generate a target text for the target image based on the sentiment label, object label, and scene label.
[0024] In the embodiment of the present invention, the server 200 can be an independent server or a server network or server cluster composed of servers. For example, the server 200 described in the embodiment of the present invention includes but is not limited to a computer, a network host, a single network server, a set of multiple network servers, or a cloud server composed of multiple servers. The cloud server is composed of a large number of computers or network servers based on cloud computing. In the embodiment of the present invention, communication between the server and the client can be achieved through any communication method, including but not limited to mobile communications based on the 3rd Generation Partnership Project (3GPP), Long Term Evolution (LTE), Worldwide Interoperability for Microwave Access (WiMAX), or computer network communications based on the TCP / IP Protocol Suite (TCP / IP) and User Datagram Protocol (UDP).
[0025] It is understood that the client 100 used in the embodiments of the present invention can be understood as a client device, which includes both receiving and transmitting hardware, i.e., a device having receiving and transmitting hardware capable of performing two-way communication over a two-way communication link. Such client devices may include: cellular or other communication devices with single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays. Specifically, the client 100 can be a desktop terminal or a mobile terminal, and can be a mobile phone, tablet computer, laptop computer, etc.
[0026] Those skilled in the art will understand that Figure 1 The application environment shown in the figure is only one application scenario of the present application solution and does not constitute a limitation on the application scenario of the present application solution. Other application environments may also include Figure 1More or fewer servers, or server network connection relationships, as shown in Figure 1 Only one server and two clients are shown. It is understandable that the copywriting generation system may also include one or more other servers, or / and one or more clients connected to the server network, which is not specifically limited here.
[0027] In some embodiments of the present invention, the work platform may be an enterprise office platform, such as enterprise WeChat. Taking server 200 as an example, it may further include an enterprise office platform contact server, an enterprise office platform configuration management server, and a Web management server. Enterprise users or developers may access the Web management server by using a Web browsing terminal to configure the field configuration information on the enterprise office platform configuration management server, and set and store the enterprise user information of enterprise employees of the enterprise office platform on the enterprise office platform contact server.
[0028] In addition, if Figure 2 As shown, Figure 2 This is a schematic diagram of a scenario of another embodiment of the document generation system provided by an embodiment of the present invention. The document generation system may further include a storage end 300 for storing data, such as a storage object database. The object database stores object data. The object data may include application templates (such as approval templates, clock-in templates, and other application templates), file data (such as Word files, Excel files, or PPT files in various formats), image data (such as images in various formats such as jpg, png, and bmp), and other types of data. Correspondingly, the object database may also be divided into multiple types of data, such as an application database, a file database, or an image database.
[0029] It should be noted that Figure 1-2 The scenario diagram of the copy generation system shown is merely an example. The copy generation system and scenario described in the embodiment of the present invention are intended to more clearly illustrate the technical solution of the embodiment of the present invention and do not constitute a limitation on the technical solution provided by the embodiment of the present invention. Those skilled in the art will appreciate that with the evolution of the copy generation system and the emergence of new business scenarios, the technical solution provided by the embodiment of the present invention is equally applicable to similar technical problems.
[0030] The following describes it in detail with reference to specific embodiments.
[0031] In this embodiment, the description will be made from the perspective of a document generation device, which may be specifically integrated into the server 200 .
[0032] The present invention provides a method for generating a copy. Figure 3 , Figure 3A flowchart of an embodiment of a method for generating a document according to an embodiment of the present invention includes: S301: Obtain the user's target image and user information; In a specific embodiment, the target image is an image uploaded and authorized by the user, and is the basis for generating the copy. The target image can be one or more images. The target image can also be a video. During the copy generation process, the image frame of the video is captured to obtain the required target frame as the target image. Obviously, information can also be directly extracted from the video to meet the needs of copy generation, which is not limited here.
[0033] S302: Perform object recognition and scene recognition on the target image respectively to obtain the object label and scene label of the target image; In a specific embodiment, object recognition refers to obtaining an object in a target image and generating a corresponding object label, and scene recognition refers to obtaining a scene in which the target image is located and generating a corresponding scene label.
[0034] It should be noted that object labels refer to labels composed of the names of objects in the image. Object labels can specifically be: apple, wine, flower, etc.
[0035] Scene labels refer to the scene reflected in the image content. Scene labels include natural landscapes, weather and climate phenomena, urban and architectural scenes, and cultural activity scenes. Specifically, they can be: mountains, rivers, lakes, and seas (an ancient town shrouded in morning mist, a snow-capped mountain top, and a forest fresh after rain), ecological wonders (waterfalls, starry skies, rainbows, and deserts), extreme astronomical phenomena (a thunderstorm, a snow-capped mountain, and a foggy morning), seasonal characteristics (a hundred flowers blooming in a spring garden), and so on. , the falling leaves on the autumn boulevard, the vast silence of the winter snowfield), modern cities (the bustling night scenes with flashing neon lights, the modern blocks with high-rise buildings), historical culture (the deep and mysterious ancient castles, the ancient streets with bluestone pavements, the urban texture reflected in the rain), celebration ceremonies (festival gatherings with singing and dancing, family reunions with bright lights), fragments of life (scientific research scenes in the laboratory, reading scenes in the library, farming scenes in the fields), etc., will not be elaborated here.
[0036] In the process of object recognition, objects in the target image can be directly identified and labeled through existing software or services.
[0037] During the scene recognition process, pre-set AI-based cloud-based image recognition and analysis services can be used to automatically detect and label scenes in target images based on deep learning technology. Obviously, multi-dimensional object recognition and scene recognition can also be performed on target images simultaneously as needed, without any restrictions here.
[0038] S303: Perform sentiment analysis on the object labels and scene labels to obtain sentiment analysis results, and modify the sentiment analysis results according to user information to obtain sentiment labels for the target image; The Emotion Wheel theory, a psychological model proposed by American psychologist Robert Plutchik in 1980, aims to systematically analyze the classification and interaction of human emotions. According to the Emotion Wheel theory, user emotions can be categorized into one or more of eight basic emotions (happiness, trust, fear, surprise, sadness, disgust, anger, and anticipation). In other words, all user emotions can be described using these limited eight basic emotions.
[0039] It should be noted that emotional labels refer to the emotional tendencies reflected by objects and scenes in pictures. The emotional tendencies are defined in words and output in a labeled manner to obtain the emotional labels reflected in the pictures. Among them, emotional labels are also divided into 8 types, namely happiness, trust, fear, surprise, sadness, disgust, anger and expectation.
[0040] In one specific embodiment, each specific object label / scene label corresponds to at least one emotion label. For example, wine (object label) can represent disgust / sadness / happiness, and sunset (scene label) can represent joy / anticipation. Therefore, to accurately obtain the minimum possible emotion labels for a target image, a global emotion analysis is performed on all object labels and scene labels of the target image to obtain a complete emotion analysis result. Furthermore, since the total number of labels corresponding to the emotion analysis results may be large, it is difficult to uniquely or minimize the emotion labels for the target image. Therefore, the emotion analysis results need to be modified based on user information to obtain a targeted emotion label for the target image that matches the user information.
[0041] S304: Generate a target text for the target image based on the emotion label, object label, and scene label.
[0042] In a specific embodiment, multimodal AI technology is used to automatically generate text based on emotion tags, object tags, and scene tags, wherein the multimodal AI technology includes but is not limited to OpenAI GPT series GPT-4 Turbo, Volcano Engine's DeepSeek-V3, deep synthesis technology (digital human + scene rendering), Z-Minds large model, diffusion model (StableDiffusion), Transformer architecture, CLIP model, etc.
[0043] In this embodiment, the object and scene in the user's target picture are identified to obtain the object label and scene label of the target picture; the emotional label of the target picture is obtained by performing sentiment analysis on the object label and scene label, and correcting the result of the sentiment analysis according to the user information, so as to automatically identify the emotional tendency of the user in sending the target picture; the target copy of the target picture is generated according to the emotional label, object label and scene label, thereby realizing the automatic generation of the target copy that meets the user's emotional needs based on the picture; since the time to obtain the user's target picture is very short and the target copy has a high degree of consistency with the user's emotions, it can not only better meet the user's needs, but also quickly obtain the target copy that meets the requirements, thereby increasing the user's enthusiasm for sending social copy and improving their social activity.
[0044] In a specific embodiment, in S301, the copy generation system can also output the target image to the outside for identification while performing object recognition and scene recognition on the target image respectively to obtain the object label and scene label of the target image, specifically including: calling the service provider interface, performing object recognition and scene recognition on the target image based on the target service provider, and obtaining the object label and scene label of the target image in the target service provider.
[0045] In this embodiment, the service provider performs object recognition and scene recognition on the image, which can effectively utilize external resources for complex data processing.
[0046] In a specific embodiment, in S303, the user information includes a user preference database, wherein the preference database is generally a data set obtained by parsing the user's historical data. In order to perform sentiment analysis on the object labels and scene labels, obtain the sentiment analysis results, and modify the sentiment analysis results based on the user information to obtain the sentiment label of the target image, specifically including: Obtain the upload time of the target image; based on the correspondence between preset tags and emotions, obtain the object emotion tag of the object tag and the scene emotion tag of the scene tag; determine the time emotion tag of the user at the upload time according to the preference database; perform emotion value frequency statistics on the object emotion tag, scene emotion tag and time emotion tag to obtain the emotion value statistics of the target image; determine the emotion tag of the target image according to the emotion value statistics.
[0047] It should be noted that the upload time of the target image can be directly obtained based on the upload attribute information of the image.
[0048] The correspondence between preset labels and emotions is a one-to-one (or multiple) correspondence between objects and emotions, or a one-to-one (or multiple) correspondence between scenes and emotions, derived by the big data system through extensive data analysis. It can also be a specially configured correspondence between objects / scenes and emotions, without limitation. Thus, for every object, there is a corresponding object label and object emotion label; for every scene, there is a corresponding scene label and scene emotion label.
[0049] In a specific embodiment, in order to quickly identify the object emotion label of the object label and the scene emotion label of the scene label, not only can the big data be parsed and semantic analysis can be directly performed on a certain object or scene to determine the corresponding emotion it represents, but also an emotion relationship library can be constructed to collect a large amount of data and determine and record the emotions that each object and scene may correspond to, so that the corresponding data can be directly captured when identifying the object emotion label and the scene emotion label, reducing time consumption.
[0050] Among them, the emotional relationship library corresponding to the object label and the represented object emotional label can be:
[0051] The emotional relationship library corresponding to the scene label and the scene emotional label represented can be:
[0052] According to the "object emotion relationship library" and the "scene emotion relationship library", after determining any object label and / or scene label in the picture, its corresponding object emotion label and / or scene emotion label can be determined quickly and accurately, which will not be elaborated here.
[0053] The preference database includes the correspondence between the user's time and the emotion tag, that is, for any specific time, there is an emotion tag determined by the user corresponding to it in the preference database, which may be one emotion tag or multiple emotion tags.
[0054] Furthermore, since object emotion labels, scene emotion labels and time emotion labels may have a one-to-one (many) relationship, in order to uniquely determine the emotion label of the target image, the frequency statistics of emotion values of object emotion labels, scene emotion labels and time emotion labels are performed, that is, each emotion label in object emotion labels, scene emotion labels and time emotion labels is exhaustively enumerated, and the frequency of occurrence of all emotion labels is counted to obtain the emotion value of each emotion label. Generally, the emotion label corresponding to the maximum emotion value is determined as the emotion label of the target image. Multiple emotion labels can also be selected as the emotion label of the target image, which will not be elaborated here.
[0055] In this embodiment, by fusing the upload time of the target image with the user's preference database, the user's time emotion label corresponding to the upload time is obtained; further, by fusing the object emotion label, scene emotion label and time emotion label, and determining the emotion label of the target image through emotion value frequency statistics, the emotional tendency of the target image is corrected by the upload time, so that the emotion label of the target image can better meet the needs of the user.
[0056] Furthermore, the preference database also includes historical data records of users. Since object emotion labels, scene emotion labels, and time emotion labels may have a certain degree of randomness, in order to obtain stable emotion value statistics of the target image, the following steps are required: Perform emotion value frequency statistics on object emotion labels, scene emotion labels and time emotion labels to obtain the initial emotion value statistics of the target image; obtain the high-frequency emotion labels with the highest frequency in the preference database, as well as the high-frequency emotion frequencies of the high-frequency emotion labels; determine the correction frequency based on the preset correction weight and the high-frequency emotion frequency; correct one or more emotions with the highest frequency in the initial emotion value statistics according to the correction frequency to obtain the emotion value statistics of the target image.
[0057] It should be noted that high-frequency emotion tags refer to emotion tags that appear in the preference database and are used most frequently by users. The high-frequency emotion tags can be one or more, and there is no limitation here.
[0058] In a specific embodiment, the preset correction weight can be set to 0.2, and the emotion frequency statistics table obtained by performing frequency statistics on the emotion tags that appear in the preference database (user's historical data) is as follows:
[0059] Then, the emotional label with the highest frequency is happiness, and the frequency of happiness is 10, that is, the high-frequency emotional frequency is 10. The preset correction weight is multiplied by the high-frequency emotional frequency to obtain a correction frequency of 2. Then, in the initial emotional value statistical result, the emotional frequency corresponding to the emotional label happiness is added by 2, that is, the emotional frequency of the high-frequency emotional label in the preference database is increased, so that the emotional value statistical result can better bias towards the user's habits. Obviously, the preset correction weight can also be other values, which are not limited here.
[0060] Obviously, the number of high-frequency emotion labels can also be adjusted, that is, the two or more emotion labels with the highest frequency can be used as high-frequency emotion labels. According to the above emotion frequency statistics table, happiness, surprise and expectation can be used as high-frequency emotion labels at the same time. Since the frequencies of surprise and expectation are both 9, the corrected frequencies of surprise and expectation are both 9*0.2=1.8. Obviously, surprise and expectation can also be corrected based on the corrected frequency 2 of the highest-frequency emotion label (happiness). I will not go into details here.
[0061] Corresponding to the high-frequency emotion labels, low-frequency emotion labels may also be set so that the emotion value statistics can better deviate from the low-frequency emotion labels.
[0062] In this embodiment, by performing emotion frequency statistics on the user's preference database, the user's most frequently appearing emotion tag is obtained, and the user's emotion tag result is corrected with the high-frequency emotion tag, thereby extracting the emotional tendency that the user most wants to show based on historical data records to better fit the user's habits.
[0063] Furthermore, the purpose of generating copy is generally to send it to a social platform, and the social platform generally allows users to set the target copy to be sent. Therefore, user information also includes the target release time of the target copy. In order to correct the results of sentiment analysis according to the target release time and obtain the sentiment label of the target image, it is specifically necessary to: obtain the time sentiment label of the target release time in the preference database; determine the time sentiment label as the user's sentiment label.
[0064] It should be noted that the preference database contains the user's historical information records. Through the historical information records, the emotional labels that the user may correspond to at certain specific time points can be obtained. That is, the time emotional label corresponding to the target release time can be directly determined according to the target release time. Since the target release time may carry a certain emotion of the user, the time emotional label is directly determined as the user's emotional label.
[0065] In a specific embodiment, the target publishing time also includes the time period in which the target publishing time falls, such as dividing 24 hours a day into multiple time periods evenly or as needed, and each time period represents one or more emotions, such as 6-10 am represents happiness / expectation, etc.
[0066] Furthermore, in order to subdivide the target release time and thus improve the reliability of the time emotion label, first, determine whether the target release time is during the preset user's birthday period; if so, select positive emotion as the time emotion label; if not, determine whether the target release time is during the preset holiday period; when the target release time is a holiday, determine the time emotion label corresponding to the holiday representative based on the preset correspondence between holidays and emotion labels; when the target release time is not a holiday, determine the time emotion label based on the time period in which the target release time is located.
[0067] Among them, the time period is obtained by dividing a day into multiple time intervals, and the time emotion labels corresponding to each time period are not completely consistent.
[0068] It should be noted that positive emotions during birthdays are generally happiness and anticipation. The correspondence between preset holidays and emotional labels can be determined by semantic analysis of big data to determine the different emotions corresponding to each holiday. The emotional relationship library corresponding to holidays and the time emotional labels they represent can be:
[0069] The emotional relationship library corresponding to different time periods and the time emotional labels represented can be:
[0070] In this embodiment, birthdays and holidays are prioritized for targeted processing. Since the emotions represented during birthdays may be relatively stable or even specific, and each holiday corresponds to different emotions due to its own different meanings, by prioritizing birthday and holiday judgments, the time emotion label corresponding to the target release time can be clearly and accurately determined; in addition, different time periods may also correspond to different time emotion labels. Therefore, in order to adaptively sort the emotion labels, for target release times that are neither birthdays nor holidays, the time emotion label is determined according to the time period in which the target release time is located, and the emotion tendency information can be adaptively extracted.
[0071] In a specific embodiment, at least one of the emotion tag, object tag, and scene tag includes multiple tags. After obtaining the emotion tag, object tag, and scene tag, in S304, in order to obtain the copy, it is necessary to generate a target copy for the target image based on the emotion tag, object tag, and scene tag, specifically including: Emotional tags, object tags, and scene tags are merged into a tag set; an object tag, a scene tag, and an emotional tag are randomly selected from the tag set and combined to obtain multiple text tag sets; multiple texts for the target image are generated based on the multiple text tag sets; and a target text for the target image is selected from the multiple texts.
[0072] It should be noted that the label set refers to all object labels and scene labels identified from the target image, as well as the emotional labels that may represent the user's current emotions or emotions when sending the target text, which are determined by combining object labels, scene labels and user information. Among them, the emotional label can include one or more, which is not limited here.
[0073] Furthermore, since the label set contains a large number of labels, and only one object label, one scene label and one emotion label are needed when generating copy, multiple copy label sets can be obtained by combining and arranging the label sets. Each copy label set corresponds to one copy, so multiple copy can be obtained based on the label set.
[0074] In order to display the target copy to the user intuitively and concisely, it is generally preferred to randomly select one from multiple copies, or all of them can be displayed at once, which is not limited here.
[0075] In a specific embodiment, the user information includes the user's spoken language data; after generating the target text for the target image based on the emotion tag, object tag, and scene tag, the following is further included: According to the colloquial data, the target copy is converted into colloquial text to obtain a colloquial copy; The colloquial data includes at least one of dialects, catchphrases, internet slang, slang and idioms.
[0076] In this embodiment, by converting the target copy into colloquial language, the colloquial copy is made more consistent with the user's language habits, which brings the relationship between the user and the copy generation system closer and improves the stickiness between the user and the social software.
[0077] It should be noted that the target text to be converted into colloquial language may be one or more, that is, all generated texts may be converted into colloquial language, and there is no limitation here.
[0078] Furthermore, after generating multiple copywritings, in order to meet the diverse needs of users, information interaction with the user can be carried out, allowing the user to actively select the final target copywriting. Specifically, all copywritings can be displayed at once, and then the user can select the target copywriting from them. Alternatively, one or more copywritings can be displayed at once according to the user's browsing habits, and the user can independently browse and query the remaining copywritings and select the final target copywriting as needed. The interaction method is not limited here.
[0079] In order to better apply the copywriting generation method provided by the embodiment of the present invention to social software, please refer to Figure 4 , Figure 4The present invention provides a flow chart of an embodiment of the social software for automatically generating copy based on the copy generation method, wherein, after logging into the social software, the user first uploads a picture / video. On the one hand, it is necessary to activate the copy generation function, and on the other hand, the uploaded picture / video needs to be extracted to determine the target picture to be used in the end; then, the target picture is subjected to image semantic analysis to identify the object label and scene label of the target picture respectively; next, the target picture is subjected to sentiment analysis based on user information, object label and scene label to obtain sentiment value statistics, and the sentiment label of the target picture is determined based on the sentiment value statistics; finally, the copy is generated based on the object label, scene label and sentiment label.
[0080] It should be noted that the part in the dotted box is the process that the copy generation device needs to handle. The process of generating copy through "object labels, scene labels and emotion labels" can also be achieved through information interaction with external service providers to directly obtain external copy data through "object labels, scene labels and emotion labels". There is no restriction here.
[0081] In addition, when the copy generation device is integrated with the social software, the user is required to independently select a unique target copy and publish it to the social software.
[0082] To facilitate better implementation of the text generation method provided in the embodiment of the present invention, the embodiment of the present invention also provides a device based on the text generation method. The meanings of the terms are the same as those in the text generation method, and the specific implementation details can be referred to the description in the method embodiment.
[0083] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of an embodiment of a document generation device provided by an embodiment of the present invention, wherein the document generation device 500 may include: Data acquisition module 501, used to obtain the user's target image and user information; The label recognition module 502 is used to perform object recognition and scene recognition on the target image to obtain the object label and scene label of the target image; The sentiment analysis module 503 is used to perform sentiment analysis on the object labels and scene labels to obtain sentiment analysis results, and to modify the sentiment analysis results according to user information to obtain the sentiment label of the target image; The text generation module 504 is used to generate a target text for a target image based on the emotion tag, object tag, and scene tag.
[0084] An embodiment of the present invention further provides an electronic device, such as Figure 6 As shown, Figure 61 is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Specifically: The electronic device may include one or more processing core processors 601, one or more computer-readable storage media memories 602, a power supply 603, an input unit 604 and other components. Those skilled in the art will understand that Figure 6 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently. The processor 601 is the control center of the electronic device, connecting the various parts of the entire electronic device using various interfaces and lines. By running or executing software programs and / or modules stored in the memory 602 and accessing data stored in the memory 602, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operation of storage media, user interface, and application programs, while the modem processor mainly handles wireless communications. It is understood that the modem processor may not be integrated into the processor 601.
[0085] Memory 602 can be used to store software programs and modules. Processor 601 executes various functional applications and data processing by running the software programs and modules stored in memory 602. Memory 602 may primarily include a program storage area and a data storage area. The program storage area may store operating storage media and applications required for at least one function (such as sound playback or image playback); the data storage area may store data generated based on the use of the electronic device. Memory 602 may also include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, memory 602 may also include a memory controller to provide processor 601 with access to memory 602.
[0086] The electronic device also includes a power supply 603 for supplying power to various components. Preferably, the power supply 603 can be logically connected to the processor 601 via a power management storage medium, thereby implementing functions such as charging, discharging, and power consumption management through the power management storage medium. The power supply 603 can also include one or more DC or AC power supplies, a recharge storage medium, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.
[0087] The electronic device may further include an input unit 604, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.
[0088] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 601 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 602 according to the following instructions, and the processor 601 will run the application programs stored in the memory 602 to implement various functions as follows: Obtain the user's target image and user information; perform object recognition and scene recognition on the target image respectively to obtain the object label and scene label of the target image; perform sentiment analysis on the object label and scene label to obtain the sentiment analysis result, and correct the sentiment analysis result according to the user information to obtain the sentiment label of the target image; generate the target copy of the target image based on the sentiment label, object label and scene label.
[0089] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0090] To this end, an embodiment of the present invention provides a computer-readable storage medium having a computer program stored thereon. The computer program is loaded by a processor to execute the steps of any of the copywriting generation methods provided in the embodiments of the present invention. For example, the computer program loaded by the processor may execute the following steps: Obtain the user's target image and user information; perform object recognition and scene recognition on the target image respectively to obtain the object label and scene label of the target image; perform sentiment analysis on the object label and scene label to obtain the sentiment analysis result, and correct the sentiment analysis result according to the user information to obtain the sentiment label of the target image; generate the target copy of the target image based on the sentiment label, object label and scene label.
[0091] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.
[0092] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0093] Since the computer program stored in the computer-readable storage medium can execute the steps in any of the copywriting generation methods provided in the embodiments of the present invention, the beneficial effects that can be achieved by any of the copywriting generation methods provided in the embodiments of the present invention can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0094] The above is a detailed introduction to a copywriting generation method, device, electronic device and storage medium provided by the embodiments of the present invention. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core ideas. At the same time, for those skilled in the art, according to the ideas of the present invention, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A copywriting generation method, characterized in that: include: Get the user's target image and user information; Performing object recognition and scene recognition on the target image respectively to obtain an object label and a scene label of the target image; Performing sentiment analysis on the object label and the scene label to obtain a sentiment analysis result, and correcting the sentiment analysis result according to the user information to obtain a sentiment label for the target image; A target text for the target image is generated according to the emotion label, the object label, and the scene label.
2. The method for generating a copywriting according to claim 1, wherein: The user information includes a preference database of the user; performing sentiment analysis on the object label and the scene label to obtain a sentiment analysis result, and correcting the sentiment analysis result according to the user information to obtain a sentiment label of the target image, including: Get the upload time of the target image; Based on the correspondence between the preset labels and the emotions, obtaining the object emotion label of the object label and the scene emotion label of the scene label; Determining a time emotion tag of the user at the uploading time according to the preference database; Performing emotion value frequency statistics on the object emotion label, the scene emotion label, and the time emotion label to obtain an emotion value statistical result of the target image; Determine the emotion label of the target image according to the emotion value statistics.
3. The method for generating a copywriting according to claim 2, wherein: The performing emotion value frequency statistics on the object emotion label, the scene emotion label, and the time emotion label to obtain the emotion value statistical result of the target image includes: Performing emotion value frequency statistics on the object emotion label, the scene emotion label, and the time emotion label to obtain an initial emotion value statistical result of the target image; Obtaining the high-frequency emotion tag with the highest frequency of occurrence in the preference database, and the high-frequency emotion frequency at which the high-frequency emotion tag appears; Determining a correction frequency based on a preset correction weight and the high-frequency emotion frequency; The one or more emotions that appear most frequently in the initial emotion value statistics are corrected according to the correction frequency to obtain the emotion value statistics of the target image.
4. The method for generating a copywriting according to claim 1, wherein: The user information includes the target release time of the target copy and the user's preference database; the step of correcting the result of sentiment analysis according to the user information to obtain the sentiment label of the target image further includes: Obtaining a time emotion label of the target release time in the preference database; Determine the time emotion tag as the emotion tag of the user.
5. The method for generating a copywriting according to claim 4, wherein: The target release time also includes the time period in which the target release time is located; and obtaining the time emotion tag of the target release time in the preference database further includes: Determining whether the target publishing time is within a preset birthday period of the user; If so, positive emotion is selected as the temporal emotion label; If not, determining whether the target release time is during a pre-set holiday; When the target release time is the holiday, determining the time emotion tag corresponding to the holiday according to the preset correspondence between holidays and emotion tags; When the target publishing time is not the holiday, determining the time emotion label according to the time period of the target publishing time; The time period is obtained by dividing a day into multiple time intervals, and the time emotion labels corresponding to each time period are not completely consistent.
6. The method for generating a copywriting according to claim 1, wherein: At least one of the emotion label, the object label, and the scene label includes a plurality of labels; Generating a target text for the target image according to the emotion label, the object label, and the scene label includes: Merging the emotion label, the object label, and the scene label into a label set; Randomly selecting one of the object labels, one of the scene labels, and one of the emotion labels from the label set and combining them to obtain multiple text label sets; Generating multiple texts for the target image according to the multiple text tag sets; A target copy for the target image is selected from the multiple copies.
7. The method for generating a copywriting according to claim 1, wherein: The user information includes the user's spoken language data; After generating a target text for the target image according to the emotion label, the object label, and the scene label, the method further includes: Performing colloquial conversion on the target text according to the colloquial data to obtain a colloquial text; The colloquial data includes at least one of dialects, catchphrases, internet slang, slang and idiomatic expressions.
8. A copywriting generating device, characterized in that: include: Data acquisition module, used to obtain the user's target image and user information; a label recognition module, configured to perform object recognition and scene recognition on the target image, respectively, to obtain an object label and a scene label of the target image; A sentiment analysis module is used to perform sentiment analysis on the object label and the scene label to obtain a sentiment analysis result, and to modify the sentiment analysis result according to the user information to obtain a sentiment label for the target image; A text generation module is used to generate a target text for the target image based on the emotion label, the object label and the scene label.
9. An electronic device, characterized in that: The system comprises a memory and a processor, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor performs the following steps: Get the user's target image and user information; Performing object recognition and scene recognition on the target image respectively to obtain an object label and a scene label of the target image; Performing sentiment analysis on the object label and the scene label to obtain a sentiment analysis result, and correcting the sentiment analysis result according to the user information to obtain a sentiment label for the target image; A target text for the target image is generated according to the emotion label, the object label, and the scene label.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program is loaded by a processor to execute the steps in the copywriting generation method according to any one of claims 1 to 10.