Hi there! What will be the consequence if an enterprise search solution that used to perform well in the past suddenly started falling short of meeting today’s needs?
This is the issue most CloudSearch users of Amazon are facing. Although CloudSearch continues to meet the needs of existing customers, AWS terminated all future access to the software as of July 2024 and no longer offers any updates. In turn, Amazon OpenSearch Serverless provides such features as semantic search, hybrid search, Retrieval-Augmented Generation (RAG), and agentic search with autoscaling.
But this migration is definitely not as simple as a switch from one AWS endpoint to another.
There are several aspects of work involved in migrating from CloudSearch to OpenSearch Serverless, including moving the data, creating field mappings, translating search queries, recreating security policies, checking the quality of the searches, and redirecting the application traffic. I will now explain the steps involved in this process in detail.
Reasons to Move from Amazon CloudSearch to OpenSearch Serverless
Amazon CloudSearch simplifies searches with the automatic setup of most of the required infrastructure. This includes support for filtering, full-text search, faceting, autocomplete, highlighting, and other standard capabilities.
Nonetheless, modern applications require semantic search, vector search, and RAG functionalities. OpenSearch Serverless provides support for such use cases along with automatic scaling of search and indexing capabilities.
This migration is hence not just about moving platforms but is an opportunity to reconsider your search, ranking, security, and retrieval processes within the context of the application’s needs.
But you must understand the cloud migration risks prior to taking any decision on it.
Step-by-Step Methods to Migrate from Amazon CloudSearch to Amazon OpenSearch Serverless
Step 1: Assess Your Current CloudSearch Environment
Before migrating to OpenSearch Serverless, assess your current CloudSearch environment. An assessment helps you identify what to migrate and what features need to be preserved.
Document Your Search Configuration
Begin with documenting the CloudSearch domain configuration, such as the partition count, instance type, replication count, number of documents, data size, and field configuration.
Note down your synonyms, analyzers, rank expressions, query behaviour, and stopwords. Know the version ofthe CloudSearch API (2011 or 2013). These versions have different capabilities.
This will allow you to migrate more than documents. You will be able to preserve the important functionality.
Determine if Serverless is Right for You
OpenSearch Serverless can be good for most CloudSearch use cases. However, not all requirements can be met using OpenSearch Serverless. For example, your application might require very consistent query latency or tight control over performance and infrastructure.
Consider these factors before the migration to prevent any redesign of the architecture afterwards.
Requirement | CloudSearch | OpenSearch Serverless |
Automatic Scaling | Yes | Yes |
Managed infrastructure | Yes | Yes |
Vector and semantic search | Limited | Yes |
Modern vector search | Limited | Yes |
Advanced OpenSearch ecosystem | No | Yes |
Manual instance control | Limited | No |
The ecosystem of AWS-DuckLabs cloud may serve as a handy broader reference for considering cloud data architecture, but your decision on migration should be based on the specific workloads rather than technology trends.
Step 2: Build the OpenSearch Serverless Collection
With the workload analysis complete, you are ready to build the destination collection.
In contrast to traditional OpenSearch domains, OpenSearch Serverless utilizes collections. It shows the principles of serverless architecture. In addition, you have to configure security layers prior to starting to move production data.
Make Index Mappings
Another important aspect is different data modeling.
Field types in CloudSearch don’t automatically translate into OpenSearch index mappings. You have to define the mappings explicitly, which will tell OpenSearch what to do with particular fields.
Thus, usually, any text field in CloudSearch may map to the text field in OpenSearch. A literal field may map to keyword, while longitude and latitude fields may be mapped to geo_point.
I would suggest considering this process more as design than a mechanical mapping process.
Step 3: Migrating and Converting the Data
Here comes the technical part of the migration.
CloudSearch doesn’t offer built-in backups or snapshots for this route; hence, keep your original data on a reliable storage service like DynamoDB or Amazon S3.
CloudSearch Document Conversion
CloudSearch relies on the Source Data Format (SDF) and needs documents that match the target mapping in OpenSearch. Thus, it's necessary to convert the data first.
Ensure that the date field is standardized, the number field is formatted, unneeded columns are deleted and renamed, JSON files are created as needed, and uploaded to Amazon S3.
OpenSearch Ingestion
For large-scale migration processes, Amazon OpenSearch Ingestion helps to transfer data from S3 to OpenSearch Serverless. This service also allows filtering, transformation, enrichment, and normalization of documents without managing any infrastructure related to ingestion. If your data ingestion is not successful, you can fix the conversion and run the pipeline once again.
So, depending on whether you require full-text search or exact match/filtered aggregation, you have to decide whether you need text or keyword mapping. Choosing the wrong mappings might lead to ambiguous search results later.
Step 4: Translating Search Queries
But moving the documents is not enough. Your app needs to know how to search them as well.
CloudSearch works on URL-based query syntax, whereas OpenSearch operates on Query DSL. Do not assume that your search queries are going to return the same results. You need to translate them instead.
For instance, an AND condition becomes a Boolean must query. Wildcards could go into the wildcard keyword, range filters into numeric_range, date_range, etc. Sort queries can be translated into the OpenSearch sort syntax.
Recreating the Effect of Boosting
In case your CloudSearch query boosts certain fields or terms, translate it into OpenSearch boosting as well; otherwise, users will get different results on the same documents.
The AWS – DuckLabs conversation is a completely different issue. For now, you need to focus on the search functionality.
Step 5: Rebuild The Security
Security is considered an aspect where you cannot bring CloudSearch concepts into OpenSearch Serverless.
OpenSearch Serverless has several levels of security policies. These should be part of your broader AWS security best practices. There are encryption policies and network policies for collection access control, secure data at rest, and data access policies that effectively define what IAM identities can execute (read, write, create indexes).
Verify thet Permissions Before Cutover
Do not find out about the permission problems at cutover time. Create a test identity representing your application and test its access rights.
Check the read and write operations prior to migrating. That's a trivial step, but it could save you from the troubles on cutover day.
That's the same principle that applies to the AWS-DuckLabs ecosystem at large: access control boundaries and clear identity are becoming essential as applications integrate several data platforms and cloud services.
Step 6: Validate Prior to Cutover
Never assume that ingestion alone means the migration has been successful. Your new collection might have all the documents but produce different results.
Do Not Forget to Verify 5 Things
The first thing is to compare the document counts for your OpenSearch and CloudSearch Serverless.
Second, validate the most significant search queries for your application and compare ranking; this step is especially relevant if field weighting, custom ranking, or boosting was implemented.
Third, test the application and latency end to end under real load conditions.
AWS recommends validating queries, document counts, latency, and ranking prior to cutover. In case of a bigger workload, representative synthetic queries will give much more information about performance than a few test cases.
A migration pilot is another step to perform. It will uncover all possible transformation, mapping, authentication, and throughput problems.
Step 7: Transition to OpenSearch Serverless
After validation is successful, make changes to your application.
Your application needs to use OpenSearch Serverless endpoints and OpenSearch Client Libraries or equivalent REST API calls.
Do not proceed immediately to shut down CloudSearch.
Maintain your original data source in case you need a fallback option when things do not behave in production as they did during testing.
Monitor your search performance, errors, query usage, and application logs during the first production window.
If all goes well, you can finally decommission the old CloudSearch environment.
What About Costs Once Migration is Complete?
With OpenSearch Serverless, it's different from the old ways of picking instances because here you have to pay for what you consume in terms of storage and compute.
It offers the ability to scale indexing and search compute independently, and the compute can be scaled to zero for idle collections, with the storage charge still continuing.
But remember that serverless doesn't equal free. Keep track of OpenSearch Compute Unit usage via Amazon CloudWatch and define reasonable limits wherever necessary. Centralized CloudWatch logs are there to enable teams to track operational and application activities.
Unexpected increases in searches will occur, thus leading to increased use of computing even though the storage remains unchanged.
Migration Errors You Need to Avoid
The first error is that of considering migration to be an easy task of transferring data. There are different data structures, queries, and security layers for CloudSearch and OpenSearch Serverless. Cloud misconfiguration may become another concern in this regard.
Do not forget about search relevancy during migration. A migration can save all documents but still deliver lower results when there is a change in ranking.
The other mistake that is made during migration is the lack of performance testing since Serverless provides automatic scaling capabilities. Though scaling makes infrastructure management easier, it is necessary to perform tests anyway.
Finally, do not delete the source data and do not roll back the migration until the production workload becomes stable.
Conclusion
Moving from Amazon CloudSearch to Amazon OpenSearch Serverless goes beyond endpoint modification.
This process entails evaluating your current configuration, building index mappings, migrating your documents, importing the documents via the ingestion pipeline, adapting CloudSearch queries for OpenSearch Query DSL, implementing new security policies, and verifying search quality prior to the migration.
In essence, what you will get is a modern search platform that can accommodate hybrid search, semantic search, vector search, and RAG without infrastructure management.
Think about the migration as an application modernization effort as opposed to a mere change of infrastructure. By safeguarding the source data, performing mappings and query testing, and verifying user behaviors, migration becomes easier.