Text Analytics processing and initialization

Reporting > Text Processing > Global Text Analytics Settings > Processings

Medallia Experience Cloud prepares comments for topic and sentiment analysis by running them through the parser. After initial processing is complete, new comments that enter Experience Cloud are processed automatically. Additionally, when you modify any of the fields in a tag pool, records associated with those fields are marked for automatic reprocessing.

The following tasks require you to initiate comment processing manually:

  • Adding a new tag pool
  • Adding a new comment field to a tag pool
  • Adding a unit group to a tag pool
  • Setting one or more languages to native processing
  • Setting one or more languages to translated processing
  • Modifying machine translation configuration
  • Modifying or creating a new user feature
  • Adding a new sentiment model
Important: Processing is aborted when deployments are in progress, so do not initiate processing during deployments. Also, do not run a Backfill while comment processing is in progress, or initiate processing when a Backfill is in progress.

Record eligibility and processing flow

To be processed for Text Analytics, a survey record must meet these requirements:

  • The record's Survey status (e_status) is COMPLETED, COMPLETION_PENDING, or EXIT_AUTO_EXCLUDED
  • There are fields filled in the survey that are in tagpools, translation, or other models
  • The survey's response date is not older than 3 years
  • The survey's response date is in the past

Text Analytics processing for new comments runs asynchronously from survey, alert, and other processes. Survey records are processed for Text Analytics in batches every 15 minutes. The batch size is configurable.

Processing speed

Comment processing is a resource-heavy operation. The length of a processing job varies based on the number of other systems running processing jobs at any given time:

  • Low-volume systems might take an hour or a few hours
  • Higher volume systems might take a day
  • Very high volume systems might take multiple days or even two weeks

Processing jobs will always parse three years worth of data regardless of tag pool configuration. The processing job must iterate through all of that to locate records that match the tag pool properties and apply themes or topics. Keep this in mind when starting a new processing job.

Medallia recommends that you do not provide hard dates to companies regarding how long processing will take. If you must communicate a time frame, communicate one longer than the pessimistic estimates listed below:

  • Optimistically — Records process at approximately 200,000 per hour.
  • Pessimistically — Records can process as slowly as approximately 50,000 per hour.
  • Translations — Translations take much longer, approximately 20,000 – 30,000 records per hour.
Tip: To save time on TA processings, run a Custom Process from Global Text Analytics Settings and narrow the time frame for the reprocessing. For example, you could process only the last year's records if analytics are not useful before that time, or an even shorter time period if you want to see a new topic introduced more recently. To save additional time, you could also narrow the conditions under which records are reprocessed using the Custom Process feature.

Processing policy

The processing policy applies to production and QA instances, and varies depending on your experience level and the number of records available. The following requirements apply to every parsing job:

  • Medallia reserves the right to stop or pause any parsing process initiated by a partner.
  • Only partners who have completed Self-Service Text Analytics training are permitted to parse.
  • If you wish to parse using Urgent Processing, request this information from your Medallia expert, who will determine the need for urgent processing and make the update (regardless of comment volume).

The following sets the specific rules you should follow:

SituationPolicy
This is the first time your instance is being parsedWork with your Medallia expert to start comment processing.
This is not the first time your instance is being parsed and it contains fewer than 50 million total recordsYou may parse on your own without contacting your Medallia expert, after completing the implementation parsing checklist below, and after validating your system setup.
This is not the first time your instance is being parsed and it contains more than 50 million total recordsYou may parse, but consider working directly with your Medallia expert, notifying them that you intend to start a processing job. Your Medallia expert needs implementation parsing checklist, below.

Processed fields

Text Analytics processes survey fields with the following attributes.

  • Fields with TEXT encoding, meaning those fields' content is plain text
  • Fields with one of the following four Content Kinds:
    • COMMENT
    • TRANSLATABLE_COMMENT
    • TEXT
    • TRANSLATABLE
Note: Comment fields are divided into phrases during parsing, and individual phrases are limited to a maximum of 5000 characters. Phrases that exceed this limit are rare but result in the record's exclusion from Text Analytics parsing, including topic tagging and sentiment analysis. In addition to the phrase character maximum, Medallia AI models process up to a maximum of 60 tokens, or approximately 60 words, per phrase. For more information, see How rules are applied.

Some fields are excluded from Text Analytics, including

  • Topic fields, such as SurveyAggregateTopicField
  • Translation fields, such as SurveyAggregateTranslationField
  • A-fields, also known as system fields
  • Multi-valued fields

The following fields are specifically excluded from Text Analytics processing.

  • EMAIL
  • PHONE
  • ADDRESS1
  • ADDRESS2
  • CITY
  • STATE
  • POSTALCODE
  • FIRSTNAME
  • LASTNAME
  • ROOM_NUMBER
  • AUTOINDEX_TEXT

Implementation parsing checklist

Use the following list before initiating text processing on your instance. Work with your Medallia expert if you need assistance.

  • Determine which Processing option is needed, as described below.
  • Note your Experience Cloud instance URL.
  • Note the number of records in your instance. Parsing a large number (more than 50 million) of records can take several days to complete.
  • If you are a Medallia partner, note your point of contact on the Medallia Partner Solutions team.
  • Note the user features created on your instance. For more information, see User features.
  • Note the unit groups and Q-fields selected in tag pools. For more information, see Creating a tag pool.
  • Check whether any deployments are running. On any Setup screen, if you see a red banner warning you not to make updates, do not initiate text processing.

Initiating text processing

Important: Do not run a Backfill while Text Analytics processing is in progress. Running a backfill and Text Analytics processing simultaneously can cause Experience Cloud downtime. For more information about running backfills, see Backfill.
  1. Prepare for comment processing, as described in Tag pool settings.

  2. Open the Reporting > Text Processing > Global Text Analytics Settings > Processings screen.

  3. Select an option from the Processing option dropdown.

  4. Click Save.

    A button appears, labeled with the name of the selected processing option.

  5. Select Confirm and then click the button to begin processing.

    Important: Parsing is aborted during the deployments. If this happens, check the deployment updates and resume parsing after the deployment ends. To resume parsing, click the button again and parsing will resume. 
  6. After comment processing begins, check the progress of the job by refreshing the page. Progress is listed in the Current or last job status table and the Last parsing executions table. When parsing is complete, the status changes from Running to Complete in both tables.

    Click Request cancel for current task if you need to stop or pause processing.

Properties

Amount of surveys in Tagging Needed (on Slug)
Indicates the number of records in need of tagging in the in-memory database.
Amount of surveys in Tagging Needed (on DB)
Indicates the number of records in need of tagging in the Experience Cloud database.
Num Tagging Failed
Indicates the number of records for which tagging failed.
Unprocessed changes
Changes Experience Cloud applies to incoming records. These configuration changes have not yet been applied to all records, so they are considered unprocessed. For example, the presence of a user feature on this list indicates:
  • The user feature term will be identified correctly for all incoming survey records.

  • Experience Cloud has not re-processed the survey records that were added to the system before the user feature was added. As a result, it is possible that the term is incorrectly identified in those records.

Configuration Type
The type of configuration change.
Identifier
Details about the configuration change. For example, if a user feature was on the list, the name of the user feature would appear.
Last finished executions

Includes this information:

Id
ID of the parsing request.
Account
Account that started parsing.
Requested
Timestamp of parsing submission.
(re)Started
If restarted, the timestamp will change to indicate the time processing resumed.
Operation
Describes the parsing request.
# Records
Number of survey records processed.
Status
Indicates the completion status. For a status of Failed, look for the following error messages:
  • ERROR: Failed batches/ Vega timeout/pump blocked — With this type of error, attempt to reprocess failed batches.
  • PUMP_TIMEOUT: contact ENG — With this type of error, reprocessing does not resolve the problem. Contact Medallia Support for assistance.
Actionable executions
A selectable list of queued and paused executions you can process.
Processing option

The table below lists all available processing options. The options available to you depend on the processings configured on your instance. If you are unsure which option to select, contact Customer Support. Note the following information before choosing a processing option:

  • The Republish all topics processing option only associates tags with surveys, and does not itself cause topics to be available in reports. To make new or updated topics available in reports, you must also publish those topics in Topic Builder. For more information, see Topics screen.
  • Reprocessing historical data removes topic tagging from inactive topics.
  • If you need to switch processing options, pause the current process, then select and run the new process.
Important: Prior to the Summer 2020 release, the Reprocess all records option applied new tags (including language detection) to all records. However, this option could be initiated only by the Medallia Engineering team.

As of the Summer 2020 release, optimizations to Text Analytics improved overall system processing, and the Reprocess all records and Custom process options were made available on the Processings screen. As of the Summer 2020 release, reprocessing affects up to three years of historical data.

Processing Function

Description

Translates

Comments

Parses

Applies

Sentiment

Applies Tags

Respects

Deactivated Topics

Respects Not yet activated Topics

Respects Topics

in Draft

Process updated translation fields

Processes any records with the new translation fields

Yes

Yes

Yes

Yes

Yes

Yes

Yes

Process updated comment fields

Processes any records with the new comment fields

No

Yes

Yes

Yes

Yes

Yes

Yes

Process updated languages

Processes all records in a language that has been added for native processing

Yes

Yes

Yes

Yes

Yes

Yes

Yes

Process updated translation fields and languagesProcesses any records with the new translation fields, and all records in a language that has been added for native processingYesYesYesYesYesYesYes

Process updated user features

For a new user feature: Processes all records that contain that user feature

For an existing user feature: Processes all records that contained the previous version of the user features and all records that contain the new user feature

No

Yes

Yes

Yes

Yes

Yes

Yes

Reprocess all recordsProcesses all recordsYesYesYesYesYesYesYes
Reprocess failed recordsProcesses records that failed during previous processingYesYesYesYesYesYesYes

Republish all topics

Republishes all changed topics

No

No

No

Yes

Yes

No

No

Retranslate all commentsRe-translates all comment fieldsYesYesYesYesYesYesYes
Custom processRepublishes comments for Medallia AI processingYesYesYesYesYesYesYes
Note: The Reprocess all records option does not perform translation processing for comments with existing translations. To retranslate comments, use the Retranslate all comments processing option.
TargetRegex
Use this property with custom processing to reprocess comments that match the specified regular expression.
Maximum allowed start date
Maximum allowed time to process. To change this value, contact Medallia Support.
Start date
Start date for a custom processing job. This processing job differs from a full processing job in that Experience Cloud processes only a subset of records. As a result, Experience Cloud does not mark all changed configurations as complete when the job finishes (the Unprocessed changes list will not change). The date format is yyyy-MM-dd HH:mm:ss.
End date
End date for a custom processing job.
Condition
A condition that restricts the scope of data to be processed in a custom processing job. For example, e_acme_survey_type = 1. For more information about conditional expression syntax, see Conditional expressions.
Comment Fields
A comma-separated list of comment fields to be processed in a custom processing job
Process only the records missing taggings
When on, processing does not override existing Text Analytics taggings (such as topics and sentiment). Only taggings that were not already on feedback records are applied.
Custom Processing Options
When you choose Custom process as the Processing option, you can select from the following Medallia AI features:
Execute Processing section
Tip: To save time on TA processings, run a Custom Process from Global Text Analytics Settings and narrow the time frame for the reprocessing. For example, you could process only the last year's records if analytics are not useful before that time, or an even shorter time period if you want to see a new topic introduced more recently. To save additional time, you could also narrow the conditions under which records are reprocessed using the Custom Process feature.

You can also select from the following options under Custom Processing Options:

  • Publish & apply all topics
    Note: The Publish & apply all topics custom processing option only associates tags with surveys, and does not itself cause topics to be available in reports. To make new or updated topics available in reports, you must also publish those topics in Topic Builder. For more information, see Topics screen.
  • Apply themes
  • Apply sentiment
  • Translate comments
  • Apply Suggested Actions
  • Apply Attention
  • Apply Customer Effort
  • Apply Recognition
  • Apply Sensitive Data rules
  • Tag Pools
    Note: Specifying options in Tag Pools determines which records are filtered for processing. All active topics are applied to the feedback records associated with your selection. If Action Intelligence features are associated with those records, they are also processed.
Current or last job
Status of the current or last job. When the job is actively in progress, the property shows two estimates:
  • Estimated percentage of records processed so far.
    Warning: This number is based on the number of records accessed during the processing job and is often inaccurate. When Experience Cloud accesses certain records twice, those records are counted twice toward the total number of records (which can results in percentages greater than 100). When Experience Cloud does not need to access every record during a processing job, the percentage can appear inaccurately low.
  • Estimated time required to complete the processing job.
    Warning: This estimate is unreliable. The length of a processing job varies based on the number of other systems running processing jobs at any given time, so the estimated time changes accordingly.
Current topic publishing activity
Indicates the status of current publishing activities.
Override merci config
Warning: Do not change this property unless instructed to do so by a Medallia expert.
Tagging Resource Strategy
Determines where to run the processing job (whether to run the job on the same server as the instance, or on a group of servers):
  • ALWAYS_LOCAL — The processing job is run on the same server that hosts your instance. This is the default setting.
  • ASYNCHRONOUS_FARM_WHEN_SENSIBLE — When the number of records being processed simultaneously exceeds a specified threshold, the processing job is run on a group of servers.
    Warning: Do not change this property unless instructed to do so by a Medallia expert.
Tagging threshold for using farm
Maximum number of records to process simultaneously. When using the ASYNCHRONOUS_FARM_WHEN_SENSIBLE strategy, this property enables you to specify the number of records to process locally before moving to the server farm. For example, if the threshold is 10,000, and 10,001 records are added to the processing queue very quickly, all 10,001 records are processed on the farm.
Warning: Do not change this property unless instructed to do so by a Medallia expert.
Size of batches on the farm
The number of records to group together in each batch. Changing this number impacts the current job.
  • When using the ALWAYS_LOCAL option, specify a batch size from 500–5,000.

  • When using the ASYNCHRONOUS_FARM_WHEN_SENSIBLE option, specify a batch size of 250.

    Warning: Do not change this property unless instructed to do so by a Medallia expert.
Max nodes on farm
Maximum number of concurrent batches to process on a farm. For the Java Parallel Process Framework (JPPF), this property also sets the maximum number of nodes in JPPF to run in parallel on the farm. Changing this number impacts the current job.
  • When using the ALWAYS_LOCAL option, specify a batch size from 1–10.

  • When using the ASYNCHRONOUS_FARM_WHEN_SENSIBLE option, specify a batch size of 30.

Warning: Do not change this property unless instructed to do so by a Medallia expert.
Batch publishing size
The number of records to add to the database at once.
Warning: Do not change this property unless instructed to do so by a Medallia expert.
Skip Messages before waiting for the pump
Indicates how many messages are sent into a processing component (called pump) that loads records into the in-memory database (called slug), before verifying that the messages were loaded.
Warning: Do not change this property unless instructed to do so by a Medallia expert.